Glossary · term

ARC-AGI

ARC-AGI is a benchmark family for testing fluid, adaptive skill acquisition on novel abstract tasks. ARC-AGI-1 and ARC-AGI-2 use static colored-grid input-output examples, while ARC-AGI-3 uses interactive, turn-based environments in which agents must explore, infer goals, and act without task instructions. No score is, by itself, proof of artificial general intelligence.

Debate2019-11-05ExternalMaturity: 4/5

Origin and context

Chollet's 2019 paper argued that intelligence evaluation should account for the efficiency with which a system acquires new skills, not only the skills it already possesses, and introduced ARC as an experimental test. ARC Prize later formalized the ARC-AGI name and released harder versions. ARC-AGI-2 preserved the few-example static-grid format, while ARC-AGI-3, introduced in 2026, extended the family to interactive environments that test exploration, goal inference, planning, and adaptation.

Sources: s1, s2, s4

Why it matters

Many benchmarks reward factual knowledge, familiar task formats, or patterns that can appear in training data. ARC-AGI instead focuses on adapting to tasks whose rule or goal must be inferred at test time. ARC-AGI-1 and ARC-AGI-2 emphasize abstraction over static transformations; ARC-AGI-3 adds interactive exploration, planning, memory, and goal acquisition. Its influence also makes evaluation discipline important: researchers must state the benchmark version, test policy, compute or action budget, and whether a result used a model alone or a larger scaffold.

Sources: s1, s2, s3, s4

Example

In ARC-AGI-1 or ARC-AGI-2, a task may show input grids and corresponding outputs in which objects are moved, recolored, or combined according to a hidden rule. In ARC-AGI-3, a system instead interacts with a novel environment over multiple turns to discover the goal and an efficient action sequence. A valid report names the ARC-AGI version and testing conditions; comparing percentages without those details can compare different tasks or resource regimes.

Sources: s1, s2, s3, s4

How it differs

Benchmark contamination

ARC-AGI is designed around novel tasks and controlled evaluation sets, while benchmark contamination is exposure to evaluation material or close derivatives during development. ARC's design reduces some memorization routes but does not remove the need for protected tests, disclosure, and contamination analysis.

Maturity and evidence

ARC has been studied since 2019, has multiple benchmark versions, an organized evaluation program, and an expanding independent literature, supporting maturity 4. The benchmark continues to evolve, so ARC-AGI is established as a research instrument rather than frozen as one immutable task set or accepted as a complete operational definition of AGI.

Sources: s1, s2, s3, s4

Limits and open questions

The static colored-grid tasks of ARC-AGI-1 and ARC-AGI-2 and the interactive environments of ARC-AGI-3 each sample only parts of general intelligence. Scores can depend on search, test-time compute, tool scaffolding, action budgets, and evaluation rules. Version changes limit direct historical comparisons. Passing a chosen score threshold does not demonstrate broad autonomy, real-world competence, reliability, or safety, and weak performance does not measure every useful form of reasoning.

Sources: s1, s2, s3, s4

Related terms

References

Last updated: 2026-08-27

In the Skills Atlas

This term is also covered in the Skills Atlas as benchmark analysis skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as llm benchmarking skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as model evaluation skill.