Glossary · term

Jagged Frontier

The jagged frontier is the uneven, task-level boundary of an AI system's useful capability: it can perform very well on one task yet fail or reduce human performance on another task that appears similarly difficult. The frontier depends on the model, version, workflow, user, tools, and evaluation criteria. It is not a fixed list of occupations that AI can or cannot perform.

Debate2023-09-16Wave 2 · 2024Maturity: 3/5

Origin and context

On 16 September 2023, coauthor Ethan Mollick publicly explained the Jagged Frontier while presenting the associated working paper as not yet peer-reviewed. A later Harvard Business School research announcement summarized the preregistered field experiment by researchers from Harvard, Wharton, Warwick, MIT, and Boston Consulting Group involving 758 consultants. AI access improved performance on tasks designed to fall within the selected model's frontier but produced worse outcomes on a task outside it. The OECD later built multi-domain capability indicators, while Stanford's 2026 AI Index discussed jagged intelligence as a related but distinct pattern within a model's capability profile.

Sources: s4, s1, s2, s3

Why it matters

The concept challenges decisions based on a model's average score, strongest demonstration, or broad occupational label. Adoption can help on some parts of a workflow and harm others, while the boundary can move after a model or tool update. Teams therefore need evaluations at the level of consequential tasks and handoffs, not only a general claim that a role is exposed to AI. Workers also need calibration skills: recognizing which outputs require verification, when independent work is safer, and how to detect that a task has crossed the current frontier.

Sources: s1, s2, s3

Example

A model may draft a clear market overview from supplied facts but make a confident strategic recommendation when the case contains a subtle constraint it cannot reliably integrate. A team that assigns the entire workflow based on the drafting success crosses the jagged frontier without measuring it. A better design evaluates each task separately, compares assisted and unassisted performance, records model and prompt versions, introduces review where errors matter, and repeats the tests after material system changes.

Sources: s1, s2

How it differs

AGI

AGI is a proposed broad capability regime whose definitions vary. The jagged frontier is an observed pattern of uneven performance in current systems and workflows. It cautions against inferring general intelligence from a set of impressive but selective results.

Benchmark contamination

Benchmark contamination can inflate a measured result because evaluation material entered training or tuning data. Jaggedness can remain even when a benchmark is clean; it concerns variation across tasks. Both problems make single-score capability claims unreliable for deployment decisions.

Jagged intelligence

Jagged intelligence describes uneven strengths and weaknesses within a model's capability profile. The jagged frontier is the task-level boundary in a human-AI workflow where assistance improves or harms outcomes. The patterns are related, but the labels are not interchangeable.

Maturity and evidence

Maturity is rated 3. The concept has an explicit empirical origin, independent academic adoption, related international capability-measurement work, and a stable analytical purpose. It is not standardized: researchers choose different task sets and definitions of success, and every model update can relocate the boundary. The original effect sizes should not be generalized beyond the studied participants, model, and consulting tasks.

Sources: s1, s2, s3

Limits and open questions

A frontier drawn from benchmarks can miss rare failures, changing environments, tool-use errors, and differences between laboratory tasks and production work. Apparent jaggedness may also reflect weak task design, insufficient elicitation, or measurement noise. The metaphor does not explain why a model fails and does not prove that every task is unpredictable. Evaluators should report task construction, baselines, system configuration, uncertainty, and whether the result measures a model alone or a human-AI workflow.

Sources: s1, s2, s3

Related terms

References

Last updated: 2026-09-04

In the Skills Atlas

This term is also covered in the Skills Atlas as model evaluation skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as benchmark analysis skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as ai output verification skill.