Computer Use
Computer use is an agent capability in which a model perceives the state of a graphical computer interface and selects actions such as moving a pointer, clicking, typing, scrolling, or using keyboard shortcuts. A surrounding runtime captures observations, executes allowed actions, and returns the changed state to the model. The model does not directly control the operating system without that application layer and its permissions.
Origin and context
OSWorld introduced a benchmark of real tasks across web and desktop applications and showed a large gap between human and model performance at publication. Anthropic released computer use in public beta in October 2024 and explicitly described it as experimental and error-prone. OpenAI presented a Computer-Using Agent in January 2025, combining visual perception and reasoning with mouse and keyboard actions and reporting both capability and safety evaluations.
Why it matters
Computer use extends automation to tasks performed through screens rather than a dedicated application API. The same observation-and-action interface can span web pages and desktop applications, as illustrated by OSWorld. That flexibility introduces a practical trade-off: success depends on interpreting the interface correctly at each step. An action sequence that works on one screen layout is not evidence of reliable operation across every application.
Example
In an illustrative workflow, an assistant opens a conference website, navigates its programme and drafts a schedule from the sessions shown. It must inspect the result of each navigation rather than assume a click succeeded. If the task later includes sending that schedule by email, the user should check the recipient and content before transmission. This confirmation recommendation follows the external-side-effect safeguards described in OpenAI's January 2025 release, not a claim that all GUI agents enforce them.
Sources: s3
How it differs
Tool Use and Function Calling
A dedicated tool integration exposes an operation through an application interface. Computer use instead selects interactions with the graphical surface, such as clicking a button or typing into a field. Both require a runtime to execute the action; a GUI-capable model does not remove that application layer.
Maturity and evidence
Maturity is rated 3 on the evidence reviewed here: an open cross-application benchmark and separately developed Anthropic and OpenAI implementations. The cited 2024 and 2025 releases document concrete capabilities and limitations, not the latest performance of every current model. They do not establish reliable unattended execution across arbitrary interfaces.
Limits and open questions
OSWorld identifies difficulty locating GUI targets and applying operational knowledge. The cited vendor releases also describe mistakes and the risk of malicious instructions on websites. OpenAI documents confirmations before external side effects and active supervision on selected sensitive sites as mitigations in its Operator implementation. These measures reduce particular risks; they are not proof of complete protection. Evaluate task outcomes and escalation behaviour in the actual deployment.
Related terms
References
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer EnvironmentsarXiv · 2024-04-11 · class A
- Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 HaikuAnthropic · 2024-10-22 · class A
- Computer-Using AgentOpenAI · 2025-01-23 · class A
Last updated: 2026-09-05
This term is also covered in the Skills Atlas as computer use ai skill.
This term is also covered in the Skills Atlas as prompt injection defense skill.
This term is also covered in the Skills Atlas as agent evaluation skill.