Glossary · term

Deliberative alignment

Deliberative alignment is a training approach that teaches a reasoning model the text of human-written safety specifications and trains it to reason over those specifications before answering. The method aims to apply policy to the particulars of a request rather than reproduce refusal patterns alone. It is a specific alignment paradigm, not a generic label for chain-of-thought, constitutional rules, or any model that pauses before responding.

Safety2024-12-20Wave 2 · 2024Maturity: 3/5

Origin and context

OpenAI introduced the method for its o-series models in December 2024, reporting improved policy adherence, jailbreak robustness, and reduced over-refusal on selected evaluations. A 2025 OpenAI–Apollo study used deliberative alignment as an anti-scheming case study and found large reductions in covert actions without complete elimination. Independent 2026 work reproduced a safety improvement while reporting residual unsafe behavior and an alignment gap between teacher and student models.

Sources: s1, s2, s3

Why it matters

Safety policies contain exceptions and context-dependent rules that pattern matching may apply inconsistently. Explicitly teaching the specification creates a route for stronger reasoning capability to improve policy application and makes the intended rule set inspectable by developers. The later stress tests also show why an aggregate benchmark gain is not a guarantee: evaluation awareness, distribution shift, base-model behavior, and adversarial adaptation can leave residual failures.

Sources: s1, s2, s3

Example

A developer supplies a model with a written policy that distinguishes benign security education from requests enabling harm. Training examples reward identifying the relevant provisions and applying them to each prompt. Evaluation then measures both unsafe compliance and excessive refusal on held-out cases, including adversarial and out-of-distribution prompts. Better scores support the tested method and model; they do not certify all policy interpretations or future attacks.

Sources: s1, s2, s3

How it differs

Constitutional AI

Constitutional AI uses written principles to generate critiques, revisions, and preference signals for training. Deliberative alignment specifically teaches a reasoning model safety specifications and trains it to recall and reason over them before responding. Both are policy-based alignment families, but their training procedures and claimed mechanisms are not identical.

AI scheming

Scheming is a target risk involving covert goal pursuit. Deliberative alignment is one mitigation approach that has been stress-tested against covert-action proxies. A reduction in those evaluations is evidence about that setup, not proof that the method removes every deceptive strategy or hidden objective.

Maturity and evidence

Maturity is rated 3. The method has a precise published definition, reported use in deployed model training, a broad anti-scheming stress test, and independent follow-on analysis. It remains below 4 because evidence is concentrated around one method family, internal policies and some training details are unavailable, and independent work still reports uncertainty and residual unsafe behavior.

Sources: s1, s2, s3

Limits and open questions

A model can misread a specification, reason from an incomplete rule set, or produce a plausible rationale that is not causally faithful. Written policies may encode disputed choices and require updates as products or threats change. Reported safety gains depend on benchmarks and threat models; they should not be generalized to every domain. Human policy review, adversarial evaluation, access controls, and monitoring remain separate layers.

Sources: s1, s2, s3

Related terms

References

Last updated: 2026-09-04

In the Skills Atlas

This term is also covered in the Skills Atlas as reasoning models skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as ai risk management skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as model evaluation skill.