ARC-AGI
ARC-AGI is a benchmark family for testing fluid, adaptive skill acquisition on novel abstract tasks. ARC-AGI-1 and ARC-AGI-2 use static colored-grid input-output examples, while ARC-AGI-3 uses interactive, turn-based environments in which agents must explore, infer goals, and act without task instructions. No score is, by itself, proof of artificial general intelligence.
Origin and context
Chollet's 2019 paper argued that intelligence evaluation should account for the efficiency with which a system acquires new skills, not only the skills it already possesses, and introduced ARC as an experimental test. ARC Prize later formalized the ARC-AGI name and released harder versions. ARC-AGI-2 preserved the few-example static-grid format, while ARC-AGI-3, introduced in 2026, extended the family to interactive environments that test exploration, goal inference, planning, and adaptation.
Why it matters
Many benchmarks reward factual knowledge, familiar task formats, or patterns that can appear in training data. ARC-AGI instead focuses on adapting to tasks whose rule or goal must be inferred at test time. ARC-AGI-1 and ARC-AGI-2 emphasize abstraction over static transformations; ARC-AGI-3 adds interactive exploration, planning, memory, and goal acquisition. Its influence also makes evaluation discipline important: researchers must state the benchmark version, test policy, compute or action budget, and whether a result used a model alone or a larger scaffold.
Example
In ARC-AGI-1 or ARC-AGI-2, a task may show input grids and corresponding outputs in which objects are moved, recolored, or combined according to a hidden rule. In ARC-AGI-3, a system instead interacts with a novel environment over multiple turns to discover the goal and an efficient action sequence. A valid report names the ARC-AGI version and testing conditions; comparing percentages without those details can compare different tasks or resource regimes.
How it differs
Benchmark contamination
ARC-AGI is designed around novel tasks and controlled evaluation sets, while benchmark contamination is exposure to evaluation material or close derivatives during development. ARC's design reduces some memorization routes but does not remove the need for protected tests, disclosure, and contamination analysis.
Maturity and evidence
ARC has been studied since 2019, has multiple benchmark versions, an organized evaluation program, and an expanding independent literature, supporting maturity 4. The benchmark continues to evolve, so ARC-AGI is established as a research instrument rather than frozen as one immutable task set or accepted as a complete operational definition of AGI.
Limits and open questions
The static colored-grid tasks of ARC-AGI-1 and ARC-AGI-2 and the interactive environments of ARC-AGI-3 each sample only parts of general intelligence. Scores can depend on search, test-time compute, tool scaffolding, action budgets, and evaluation rules. Version changes limit direct historical comparisons. Passing a chosen score threshold does not demonstrate broad autonomy, real-world competence, reliability, or safety, and weak performance does not measure every useful form of reasoning.
Related terms
References
- On the Measure of IntelligenceFrançois Chollet / arXiv · 2019-11-05 · class A
- ARC-AGI-2: A New Challenge for Frontier AI Reasoning SystemsARC Prize Foundation / arXiv · 2025-05-17 · class A
- The ARC of Progress towards AGI: A Living Survey of Abstraction and ReasoningVahdati et al. / arXiv preprint · 2026-03-09 · class B
- ARC-AGI-3: A New Challenge for Frontier Agentic IntelligenceARC Prize Foundation / arXiv · 2026-03-24 · class A
Last updated: 2026-08-27
This term is also covered in the Skills Atlas as benchmark analysis skill.
This term is also covered in the Skills Atlas as llm benchmarking skill.
This term is also covered in the Skills Atlas as model evaluation skill.