Glossary · term

AI control

AI control is a safety research approach for using a capable but potentially untrusted model while limiting its ability to cause unacceptable outcomes, including when it may deliberately subvert oversight. It combines deployment protocols such as monitoring, trusted editing, restricted affordances, audits, and escalation, then red-teams the complete protocol against attack strategies. The defining assumption is adversarial behavior, not merely accidental model error.

Safety2023-12-12Wave 2 · 2024Maturity: 3/5

Origin and context

The research program was documented in a preprint submitted in December 2023 and later published at ICML 2024. The paper called AI control a second layer of defense: instead of only trying to make a model avoid harmful actions, it evaluated protocols intended to prevent harm even if the model tried to cause it. Its testbed used an untrusted GPT-4, a trusted GPT-3.5, limited trusted labor, and programming tasks with hidden backdoors. A 2026 ICLR paper independently showed that adaptive prompt-injection attacks could evade several monitor-based protocols, expanding the evidence without generalizing beyond evaluated settings.

Sources: s1, s2, s3

Why it matters

More autonomous systems can take sequences of actions, use tools, and produce outputs that are difficult to inspect manually. If a deployment relies only on the model being aligned, one strategic failure can bypass ordinary quality controls. AI control asks what trusted components, information boundaries, sampling rules, and human review budgets can reduce that risk while preserving useful work. It also makes the threat model and the red team's permitted knowledge part of the evaluation rather than hidden assumptions.

Sources: s1, s2

Example

A coding agent proposes repository changes, but cannot merge directly. A monitor scores suspicious behavior, high-risk actions go to a trusted editor or human, and a randomized audit samples apparently safe work. Evaluators then give an attack policy knowledge of the protocol and test whether it can insert backdoors without detection. Passing that experiment supports only the stated task, models, audit budget, and attacker capabilities; it is not proof that the protocol controls every deployment.

Sources: s1, s2

How it differs

AI Guardrails

Guardrails are broad runtime restrictions or checks for policy, security, and product behavior. A guardrail can be one component of an AI-control protocol. AI control adds an explicit untrusted-model threat model and evaluates whether the combined protocol survives intentional, adaptive subversion.

Constitutional AI

Constitutional AI is a training approach intended to shape model behavior using written principles and AI feedback. AI control is a deployment-side defense that does not assume training made the model compliant. Training and control can form separate layers in one safety strategy.

AI safety cases

A safety case is a structured argument linking claims to evidence for a defined system and context. Control evaluations can supply evidence to such an argument, but the safety case is not itself a monitor, access restriction, or control protocol.

Maturity and evidence

Maturity is rated 3. AI control has a peer-reviewed foundational formulation, concrete protocols and benchmarks, and independent peer-reviewed attack research. It remains below 4 because the evidence concentrates on limited coding environments and model pairings, adaptive attacks expose major weaknesses, and no protocol has demonstrated robust coverage of arbitrary high-capability deployments.

Sources: s1, s2

Limits and open questions

Control results depend on what the attacker knows, which actions the untrusted model can take, monitor independence, trusted-model capability, audit budget, and the cost assigned to failures. A monitor can become a single point of failure or an attack surface. Protocol performance can change as models and tools change. Claims should report both usefulness and safety under explicit threat models, and should not turn a benchmark result into an assurance of containment.

Sources: s1, s2

Related terms

References

Last updated: 2026-09-04

In the Skills Atlas

This term is also covered in the Skills Atlas as ai risk management skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as ai red teaming skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as ai guardrails skill.