AI control
AI control is a safety research approach for using a capable but potentially untrusted model while limiting its ability to cause unacceptable outcomes, including when it may deliberately subvert oversight. It combines deployment protocols such as monitoring, trusted editing, restricted affordances, audits, and escalation, then red-teams the complete protocol against attack strategies. The defining assumption is adversarial behavior, not merely accidental model error.
Origin and context
The research program was documented in a preprint submitted in December 2023 and later published at ICML 2024. The paper called AI control a second layer of defense: instead of only trying to make a model avoid harmful actions, it evaluated protocols intended to prevent harm even if the model tried to cause it. Its testbed used an untrusted GPT-4, a trusted GPT-3.5, limited trusted labor, and programming tasks with hidden backdoors. A 2026 ICLR paper independently showed that adaptive prompt-injection attacks could evade several monitor-based protocols, expanding the evidence without generalizing beyond evaluated settings.
Why it matters
More autonomous systems can take sequences of actions, use tools, and produce outputs that are difficult to inspect manually. If a deployment relies only on the model being aligned, one strategic failure can bypass ordinary quality controls. AI control asks what trusted components, information boundaries, sampling rules, and human review budgets can reduce that risk while preserving useful work. It also makes the threat model and the red team's permitted knowledge part of the evaluation rather than hidden assumptions.
Example
A coding agent proposes repository changes, but cannot merge directly. A monitor scores suspicious behavior, high-risk actions go to a trusted editor or human, and a randomized audit samples apparently safe work. Evaluators then give an attack policy knowledge of the protocol and test whether it can insert backdoors without detection. Passing that experiment supports only the stated task, models, audit budget, and attacker capabilities; it is not proof that the protocol controls every deployment.
How it differs
AI Guardrails
Guardrails are broad runtime restrictions or checks for policy, security, and product behavior. A guardrail can be one component of an AI-control protocol. AI control adds an explicit untrusted-model threat model and evaluates whether the combined protocol survives intentional, adaptive subversion.
Constitutional AI
Constitutional AI is a training approach intended to shape model behavior using written principles and AI feedback. AI control is a deployment-side defense that does not assume training made the model compliant. Training and control can form separate layers in one safety strategy.
AI safety cases
A safety case is a structured argument linking claims to evidence for a defined system and context. Control evaluations can supply evidence to such an argument, but the safety case is not itself a monitor, access restriction, or control protocol.
Maturity and evidence
Maturity is rated 3. AI control has a peer-reviewed foundational formulation, concrete protocols and benchmarks, and independent peer-reviewed attack research. It remains below 4 because the evidence concentrates on limited coding environments and model pairings, adaptive attacks expose major weaknesses, and no protocol has demonstrated robust coverage of arbitrary high-capability deployments.
Limits and open questions
Control results depend on what the attacker knows, which actions the untrusted model can take, monitor independence, trusted-model capability, audit budget, and the cost assigned to failures. A monitor can become a single point of failure or an attack surface. Protocol performance can change as models and tools change. Claims should report both usefulness and safety under explicit threat models, and should not turn a benchmark result into an assurance of containment.
Related terms
References
- AI Control: Improving Safety Despite Intentional SubversionICML / PMLR · 2024-07 · class A
- Adaptive Attacks on Trusted Monitors Subvert AI Control ProtocolsICLR · 2026 · class A
- AI Control: Improving Safety Despite Intentional SubversionarXiv; later ICML · 2023-12-12 · class A
Last updated: 2026-09-04
This term is also covered in the Skills Atlas as ai risk management skill.
This term is also covered in the Skills Atlas as ai red teaming skill.
This term is also covered in the Skills Atlas as ai guardrails skill.