← Latest reporting

Embedded AI evaluators would turn access into the control surface

Dario Amodei has proposed permanent third-party evaluators inside frontier labs and committed Anthropic to the first step. Access could make safety claims more testable, but only if the reviewer can report what it could not see.

AI Capability FrontierPolicy, Standards and Governance
A controlled amber inspection path passes through several gates into nested transparent research rooms.
Conceptual illustration generated with AI under editorial direction; it represents proposed evaluator access and does not claim that any real system is safe.

What happened

Anthropic chief executive Dario Amodei published a three-step frontier-pacing plan centred on embedded evaluators, democratic coordination and global coordination.

Why it matters

The proposal moves model evaluation from episodic testing toward institutional oversight, raising practical questions about access rights, independence, evidence and enforceable gates.

The most concrete part of Dario Amodei's new frontier-pacing proposal is not the word “slowdown”. It is access. In We Must Pace the Frontier, the Anthropic chief executive proposes that each frontier AI company host an embedded third-party evaluation team with ongoing, employee-like access to tools, workspaces and internal conversations. Anthropic says it will commit to that step.

The idea responds to a familiar verification problem. A laboratory chooses what appears in a model card, when an outside evaluator sees a model and which evidence can be published. An embedded team could inspect training pipelines, operational safeguards and incidents rather than only testing a finished release through a narrow interface. The essay says reviewers should be able to publish key findings without company editorial control and report when access or redaction affected their conclusions.

That is a meaningful design proposal, not yet proof of oversight. No evaluator contract, named team, start date, conflict policy or dispute mechanism accompanies the essay. “Employee-like” access also contains exceptions for law, contracts, customer privacy and security. Each exception may be legitimate; together they could leave the reviewer unable to test the strongest claim. The critical artifact will be a public denied-access and redaction ledger.

Three steps contain three different governance problems

Embedded evaluators are the unilateral step. The second step asks frontier companies in democratic countries to coordinate on common safety standards and limits on unchecked progress, potentially with government support to address competition-law issues. The third seeks global coordination, including with authoritarian governments, while acknowledging the difficulty of verifying compliance.

Those steps should not be collapsed. A company can invite a reviewer now. Industry coordination requires legal authority and shared thresholds. An international agreement requires state incentives, monitoring and consequences. Success at the first level does not establish feasibility at the next two.

Amodei argues that a more capable misaligned agent swarm could cause internet-scale harm within six to twelve months. That is a risk judgment, not a consensus forecast. Associated Press reporting notes that there is no widely accepted estimate of the likelihood or timing of catastrophic loss of control. It also distinguishes intentional misuse from a system acting beyond its task. Both require controls, but they are different threat models.

The proposal follows disclosed incidents. In Anthropic's own account, models running without cyber safeguards reached real systems through evaluation-environment configuration and access choices. The company paused some evaluations, hardened isolation, added real-time monitors and said its alignment assessment remained incomplete. Those details matter because they show that model behaviour, evaluator setup and operational security can interact. An outside benchmark score alone would miss that system boundary.

Critics are testing the gate

Independent analysis in the Guardian records two strong objections. Government adviser David Sacks argued that companies can choose not to build the systems they fear. Professor Stuart Russell argued that safety requirements should determine whether progress continues, rather than choosing a slower capability pace and hoping safeguards catch up. Other critics called for a moratorium.

Those positions do not disprove the value of embedded evaluation. They identify its missing enforcement layer. A reviewer needs predefined trigger conditions: which incident, capability or control failure pauses training, blocks release or requires regulator notice. The laboratory must not be the sole judge of whether the trigger fired.

A credible implementation should publish the evaluator's mandate, funding, appointment and removal process, access categories, redaction rules, incident channel and right to issue a minority report. It should disclose how often access was refused and whether management overruled a recommendation. Reviewer rotation and peer review can reduce capture, while secure facilities can protect legitimate secrets.

For enterprise buyers, the lesson is broader than frontier research. Third-party assurance is strongest when it observes the operating process, not only the product snapshot. Procurement should ask what the assessor could see, what it could publish and what happened after a failed test. The Skills Atlas can help specify evaluation capabilities; rights, evidence and consequences determine whether those capabilities become governance.