Glossary · term

Agent harness

An agent harness is the runtime scaffolding around a model that turns repeated model calls into an operating agent. It commonly manages the control loop, tool dispatch, state, context assembly, policies, errors, and output handling. Some implementations also provide memory, approvals, tracing, or handoffs. The harness is distinct from the underlying model and from the external environment in which tools execute; there is no single required component list or standardized harness interface.

Agents2025-10-06Wave 3 · 2025–26Maturity: 3/5

Origin and context

At OpenAI DevDay in October 2025, Cursor described the agent harness and tools around its coding system. Anthropic used the framing in November 2025 for infrastructure that lets coding agents work coherently across multiple context windows. OpenAI's April 2026 Agents SDK evolution and Microsoft's August 2026 Agent Framework documentation describe related orchestration and harness layers; a May 2026 research preprint separately analyzed harness construction. Together these sources show a converging engineering category, but not a single inventor or standardized boundary.

Sources: s7, s8, s1, s2, s3, s4

Why it matters

A capable model alone does not decide how tools are exposed, when state is persisted, what context survives a long task, or how failures are retried and surfaced. Harness choices shape cost, observability, reproducibility, and the boundary of agent action. They also provide places to enforce deterministic controls around a probabilistic model, such as tool allowlists, approval checkpoints, budgets, and structured traces.

Sources: s7, s1, s2, s3, s4

Example

Consider a coding assistant working across several sessions. The harness supplies tools and instructions, assembles context, records progress and makes the next model call after each tool result. A later session can recover project state from saved artifacts instead of relying on an exhausted conversation window. In this illustrative architecture, a separate sandbox executes commands. The distinction matters: changing where a command runs is not the same as changing the agent loop or its context-management policy.

Sources: s1, s2

How it differs

Agent identity

The harness manages runtime behavior and can attach credentials or identity context to actions. Agent identity represents the principal that is acting and its delegation relationships. A harness may consume an identity service, but it cannot turn a shared credential into a distinct, auditable principal merely by logging it.

Evaluation-driven development (EDD)

An evaluation harness runs test cases and scores system behavior; an agent harness runs the operational loop. One system can contain both, and production traces from the agent harness can inform evaluations, but the terms should not be treated as synonyms.

Harness Engineering

An agent harness is the runtime artifact around a model. Harness engineering is the practice of designing, testing, and iteratively improving that scaffolding in response to observed behavior. The concepts are closely related but not exact synonyms: one names the system, while the other names the engineering work performed on it.

Maturity and evidence

Maturity is rated 3 for a documented engineering category, not a universal architecture. Anthropic's Claude Agent SDK, OpenAI's Agents SDK and Microsoft's Agent Framework use the harness framing for concrete runtime responsibilities around model calls, tools and context. An independent research preprint studies the same artifact. This cross-organization implementation evidence establishes the narrow runtime meaning. It does not establish comparable performance, interchangeable interfaces or a standard list of components.

Sources: s1, s2, s3, s4

Limits and open questions

A harness does not guarantee successful long tasks. Anthropic documents incomplete work, premature completion claims and context lost between sessions; its demonstrated workflow is not evidence that the same design succeeds in every domain. OpenAI separately distinguishes the orchestration layer from the environment that executes code. Comparing harnesses therefore requires stating which tools, persistence mechanisms, permissions and recovery behavior are included, rather than treating the label as a reliability or security certification.

Sources: s1, s2

Related terms

References

Last updated: 2026-09-05

In the Skills Atlas

This term is also covered in the Skills Atlas as ai agent design skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as code execution agents skill.