Glossary · term

Constitutional Classifiers

Constitutional Classifiers are input and output classifiers trained from a written set of content rules and synthetically generated examples. They are placed around a language model to detect requests or responses that fall within specified harmful-content categories. The constitution defines the classification policy; the method is a defense layer, not a guarantee that every jailbreak or harmful output will be blocked.

Safety2025-01-31Wave 2 · 2024Maturity: 3/5

Origin and context

Anthropic's arXiv-only preprint was submitted on 31 January 2025 and its accompanying article appeared on 3 February. The authors generated training data by using language models to transform a natural-language constitution into examples, then trained separate input and output classifiers. They evaluated the prototype through human red teaming and automated attacks. A later independent arXiv-only preprint explicitly attacked Constitutional Classifiers, showing that the term and target architecture were understood outside the originating organization.

Sources: s1, s2, s3

Why it matters

A separately trained classifier can make a safety policy more explicit and can screen both what enters and what leaves a model. This creates an additional control point that teams can evaluate, update, and monitor without assuming the generative model will consistently police itself. The design also exposes practical trade-offs: policy coverage, false refusals, adaptive attacks, latency, and compute cost must be measured for the actual model, classifier thresholds, language, and traffic pattern.

Sources: s1, s2, s3

Example

A service can run an input classifier before sending a request to its model and an output classifier before returning the answer. If either classifier detects a category defined by the constitution, the service can refuse or route the exchange for review. A system prompt that merely says 'do not provide harmful advice' is not a Constitutional Classifier: it lacks the separate trained classification components and their policy-derived data pipeline.

Sources: s1, s2

How it differs

Constitutional AI

Constitutional AI is a broader alignment approach that uses principles to guide critique, revision, and preference feedback during training. Constitutional Classifiers use a constitution to train external input and output filters. They share a policy-document idea but operate at different layers and should remain separate entries.

LLM jailbreaking

Jailbreaking is the adversarial objective or technique of bypassing safeguards. Constitutional Classifiers are one proposed defensive architecture. Success against one configuration does not establish that every classifier is ineffective, while a low attack-success rate in one test does not establish universal robustness.

Maturity and evidence

Maturity is rated 3. The architecture is specified in a detailed primary arXiv preprint, was tested through multiple attack procedures, and has become a named target of an independent adversarial arXiv preprint. Neither preprint is presented as peer reviewed. The architecture remains below broad operational maturity because evidence is concentrated on a limited set of configurations, no common implementation standard exists, and adaptive-defense performance can change with the threat model.

Sources: s1, s2, s3

Limits and open questions

The reported 4.4 percent jailbreak success rate belongs to Anthropic's stated automated evaluation setup and should not be generalized to all attackers or deployments. Classifiers can miss novel attacks, over-block benign requests, inherit gaps in synthetic data, and add inference cost. Independent work has demonstrated black-box attacks against the defense. Claims should therefore state the configuration, policy scope, attack budget, baseline, and false-positive trade-off.

Sources: s1, s2, s3

Related terms

References

Last updated: 2026-09-04

In the Skills Atlas

This term is also covered in the Skills Atlas as ai guardrails skill.