Gaia2
Gaia2 is an agent benchmark built on Meta Agents Research Environments (ARE). It places an LLM-based agent in a simulated consumer environment containing apps, data and timed events. Unlike a static question set, the environment can change while the agent is working. Scenarios test state-changing execution, search, adaptation, temporal constraints, ambiguity, controlled noise and communication with simulated application agents.
Origin and context
Gaia2 first appeared with ARE in September 2025 as a successor to the read-oriented GAIA benchmark. Its dedicated paper was accepted as an ICLR 2026 oral. The public ARE package, Gaia2 dataset and leaderboard workflow make the scenarios runnable, while later work extended part of the suite across ten languages. The benchmark name does not mean a numbered release of every dataset called GAIA.
Why it matters
An agent can answer static questions yet fail when an email arrives mid-task, an API returns noise, a deadline passes or an instruction needs clarification. Gaia2 exposes those failure modes in a repeatable simulated world. Expected state-changing actions and ordering or timing constraints support action-level verification rather than scoring only the final response. The same structure can also generate verified trajectories for debugging or reinforcement learning from verifiable rewards.
Example
A team connects its agent scaffold to ARE, pins the model and provider configuration, then runs repeated Gaia2 scenarios by capability. A calendar task may require the agent to inspect existing events, ask about a conflict and perform writes before a timed event changes the state. The team compares overall success with per-capability results, cost and latency, then inspects structured traces. A valid report records the benchmark version, scaffold, judge configuration, budgets, retries and run variance.
How it differs
Agent sandboxes
ARE supplies a simulated evaluation environment. An agent sandbox isolates execution resources; it does not by itself define Gaia2 tasks, expected actions or scoring.
Reinforcement Learning with Verifiable Rewards (RLVR)
RLVR is a training approach. Gaia2's verifiers can provide rewards for training, but the benchmark can also be used only for evaluation and does not prescribe one learning algorithm.
LLM-as-a-judge
Gaia2 uses exact checks for rigid fields and an LLM rubric for some flexible text. Its write-action verification is therefore broader than, but not independent of, LLM-as-a-judge techniques.
Benchmark contamination
Benchmark contamination concerns exposure of test material. Gaia2's additional concerns include synthetic-world validity, harness dependence, judge behavior and variance from asynchronous execution.
Maturity and evidence
Maturity is 3. Gaia2 has peer-reviewed publication, open code and data, a documented runner, model comparisons, an external beta implementation and a multilingual derivative. It is not rated higher because independent score reproduction remains thin, one external implementation warns that parity is unvalidated, the benchmark is still evolving and synthetic scenarios cannot establish production reliability by themselves.
Limits and open questions
Gaia2 models a fictional app ecosystem rather than uncontrolled workplaces or the open web. Benchmark scores are joint measurements of the model, agent loop, prompts, tool descriptions, timeouts, provider latency and judge setup. Asynchronous scenarios can vary between runs, and some flexible fields use an LLM rubric. Published rankings age quickly as model endpoints change. MASEval's separate implementation is useful adoption evidence but explicitly lacks validation against original results; Benchgen had no external run results at review time. Connecting ARE to real tools or unsafe MCP servers changes the risk boundary and requires separate isolation and permissions.
Related terms
References
- ARE: Scaling Up Agent Environments and EvaluationsarXiv · 2025-09-21 · class A
- Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous EnvironmentsICLR / arXiv · 2026-02-12 · class A
- Meta Agents Research EnvironmentsMeta · 2025 · class A
- Gaia2 and ARE: Empowering the Community to Evaluate AgentsHugging Face and Meta Agents Research Environments · 2025 · class A
- GAIA2: Dynamic Multi-Step Scenario Benchmark (Beta)MASEval · 2026 · class B
- GAIA2 — BenchgenBenchgen · 2026-08-10 · class B
- OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI AgentsarXiv · 2026-08-09 · class B
Last updated: 2026-09-07
This term is also covered in the Skills Atlas as agent evaluation skill.
This term is also covered in the Skills Atlas as llm benchmarking skill.
This term is also covered in the Skills Atlas as llm evaluation design skill.
This term is also covered in the Skills Atlas as reinforcement learning from verifiable rewards skill.
This term is also covered in the Skills Atlas as multi agent coordination patterns skill.
This term is also covered in the Skills Atlas as agent sandboxing skill.