March of Nines
The March of Nines is an engineering metaphor for the repeated work needed to move a system from a plausible demo toward high reliability: from roughly 90% success to 99%, then 99.9%, and onward. In AI-agent discussions it warns that impressive behavior on selected examples is only an early milestone. Each additional nine exposes rarer conditions, integration failures and operational demands that require new evaluation and engineering.
Origin and context
Reliability engineering has long described availability targets by their number of nines. The earliest exact phrase found in this review is Elon Musk's `long march of nines` on Tesla's July 2020 earnings call about autonomous driving. In October 2025, Andrej Karpathy applied the metaphor to software agents, saying that years at Tesla made him skeptical of demos and that successive reliability levels each demanded another substantial block of work.
Why it matters
Averages conceal the production gap. A workflow may look useful while its failures cluster in unusual users, tool states or long task chains. The metaphor directs teams to define success at the end-to-end task level, measure failure severity as well as frequency, inspect long tails, and budget for recovery, monitoring and escalation. It also challenges forecasts that infer deployment speed directly from a prototype's apparent capability.
Example
A team evaluating a support agent could freeze a representative task set, record tool calls and final outcomes, separate harmless formatting mistakes from unauthorized actions, and rerun the suite after every change. Arithmetic, policy checks and state transitions that can be deterministic should leave the model path. Residual failures need validators, retry limits, observability and a human handoff instead of an unsupported claim that a benchmark percentage makes the agent production-ready.
How it differs
LLM evaluations (evals)
Evals are the tests and measurement process. The March of Nines is a metaphor for why increasingly demanding evaluation and remediation continue after an initial success rate looks high.
Agent observability
Agent observability provides traces and operational evidence needed to locate failures. It is one tool for advancing reliability, not another name for the reliability journey.
Decade of Agents
Decade of Agents is Karpathy's timeline framing. March of Nines is the separate engineering intuition he used to explain why dependable agents may take years to build.
Maturity and evidence
Maturity is 3. The phrase has traceable provenance, a clear current meaning and independent reuse across engineering publishing, technology reporting and institutional analysis. It is not rated higher because its central effort claim is heuristic, teams operationalize reliability differently, and no shared measurement standard defines which nine an AI system has reached.
Limits and open questions
The metaphor must not be read as a mathematical law. Adding a nine reduces the remaining error rate by an order of magnitude; it does not prove that labor, calendar time or cost rises by exactly 10x. Multiplying per-step probabilities is valid only under the stated model, especially independence, and can mislead when errors are correlated, retried, detected or recoverable. Availability nines measure service uptime, whereas an agent may be available yet semantically wrong. Required reliability depends on task distribution, failure consequence and human safeguards; safety-critical release decisions need domain-specific evidence and review.
Related terms
References
- Tesla Q2 2020 Earnings Call TranscriptThe Motley Fool · 2020-07-23 · class B
- Andrej Karpathy — AGI is still a decade awayDwarkesh Podcast · 2025-10-17 · class A
- Service Level ObjectivesGoogle Site Reliability Engineering · 2016 · class A
- Keep Deterministic Work DeterministicO'Reilly Radar · 2026-03-19 · class B
- AI Will (Eventually) Turbocharge Productivity and ProfitsTD Asset Management · 2026 · class B
- Karpathy's March of Nines shows why 90% AI reliability isn't even close to enoughVentureBeat · 2026-03-06 · class B
Last updated: 2026-09-07
This term is also covered in the Skills Atlas as agent evaluation skill.
This term is also covered in the Skills Atlas as llm evaluation design skill.
This term is also covered in the Skills Atlas as ai output verification skill.
This term is also covered in the Skills Atlas as ai risk management skill.