Glossary · term

Constitutional AI

Constitutional AI (CAI) is a model-alignment approach that uses an explicit set of written principles to guide critique, revision, and preference feedback. In the originating method, a model first revises responses against constitutional principles, then AI-generated preference comparisons support reinforcement learning from AI feedback. The constitution supplies supervisory criteria; it is not a legal constitution and does not remove human choices about values.

Safety2022-12-15Wave 1 · 2023Maturity: 3/5

Origin and context

Anthropic introduced Constitutional AI in a paper submitted in December 2022. The method combined supervised critique-and-revision with a reinforcement-learning phase based on AI feedback, aiming to reduce harmful responses while preserving helpfulness and making the normative basis more transparent. Independent work later proposed IterAlign, which searches for additional principles from observed model failures, illustrating both continued interest and the burden of relying on a fixed hand-written constitution.

Sources: s1, s2

Why it matters

CAI makes some behavioral criteria inspectable rather than leaving every preference implicit in a large annotation set. It can scale feedback generation and support discussion about which principles a system follows. The difficult governance questions remain: who selects the constitution, how conflicts between principles are resolved, which cultures and affected groups are represented, and whether behavior matches the text in new contexts. Transparency of principles is useful evidence, not proof of alignment.

Sources: s1, s2, s3

Example

A developer might ask a model to answer a harmful request, critique that answer against a principle prohibiting facilitation of serious harm, and produce a safer revision. Many such comparisons can train a preference model or policy. If reviewers merely add a safety prompt at inference time, they are using prompt-based guardrails, not the full CAI training method. Human governance is still required to approve principles and test their consequences.

Sources: s1, s2

How it differs

RLHF

RLHF is a broad family that learns from human preferences. Constitutional AI specifies principles and, in its reinforcement phase, uses model-generated preference feedback under human-authored supervision. The methods can share optimization machinery, but their feedback sources and governance design differ. CAI should therefore remain a separate entry linked to, not merged with, RLHF.

Maturity and evidence

Maturity is rated 3. CAI has a detailed primary method, reported experiments, independent extensions, and a stable name, so it is beyond an early proposal. Evidence remains concentrated relative to more mature training techniques, implementations vary, and there is no shared standard for choosing or auditing constitutions. Broader independent replication and governance practice would support a higher rating.

Sources: s1, s2, s3

Limits and open questions

Written principles can be incomplete, ambiguous, culturally narrow, or internally inconsistent. A model may apply them differently across prompts, languages, or adversarial settings, and AI feedback can reproduce the evaluator model's blind spots. Public principles do not reveal every training choice. Evaluations should test conflicts, over-refusal, disparate effects, and behavior outside the training distribution, with independent human review of both the constitution and outcomes.

Sources: s1, s2, s3

Related terms

References

Last updated: 2026-08-27

In the Skills Atlas

This term is also covered in the Skills Atlas as RLHF skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as reward modeling skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as ai guardrails skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as ai red teaming skill.