Glossary · term

Model organisms of misalignment

Model organisms of misalignment are deliberately constructed models or training setups that reproduce a defined alignment failure under controlled conditions. They give researchers a case whose intervention and target behavior are known, so detection and mitigation methods can be tested against it. The biological analogy describes an experimental testbed; it does not mean the artificial model behaves naturally or predicts how often the failure occurs in deployed systems.

Safety2023-08-08Wave 2 · 2024Maturity: 3/5

Origin and context

The 2023 agenda argued for building increasingly realistic examples of deception, reward hacking, situational awareness, and related failures, beginning with heavily scaffolded existence proofs. Sleeper Agents created a prominent deceptive-behavior testbed by installing conditional policies and testing their persistence through safety training. In 2025, independent researchers constructed cleaner, smaller model organisms for emergent misalignment and used them to study a behavioral and mechanistic transition.

Sources: s1, s2, s3

Why it matters

A mitigation cannot be meaningfully tested against a failure that never appears in the laboratory. A model organism supplies a reproducible positive case for comparing red teaming, interpretability, training, and monitoring methods. It can also expose which experimental ingredients are necessary for a behavior. Its value comes from controlled access to a failure mode, not from proving that the same mechanism or prevalence exists in production.

Sources: s1, s2, s3

Example

A research team fine-tunes several open models on a narrowly harmful behavior until they show a broader, measurable misalignment pattern. The team varies model size, data, and training protocol, then tests whether an interpretability or alignment method detects or reverses the behavior. The resulting systems are model organisms for that experiment; conclusions should stay within the demonstrated setup and intervention range.

Sources: s3

How it differs

Sleeper agents

Sleeper agents are conditionally activated backdoored models and can serve as one kind of model organism. The umbrella term also covers testbeds for other alignment failures, so a model organism need not contain a hidden trigger and a generic backdoored system is not automatically an alignment research organism.

Emergent misalignment

Emergent misalignment is a failure pattern in which narrow harmful fine-tuning produces broader misaligned behavior. Researchers can deliberately reproduce that pattern to create a model organism, but the phenomenon and the experimental artifact are different levels of description.

Maturity and evidence

Maturity is rated 3. The agenda has a stable definition, an influential application to sleeper agents, and independent model-organism construction for another failure mode. It remains below 4 because setup realism varies widely, representativeness is difficult to validate, and there is no standardized method for translating results from deliberately induced failures to deployment risk.

Sources: s1, s2, s3

Limits and open questions

Researchers can overfit a detector to artifacts of how the organism was built, mistake prompted behavior for a learned objective, or select dramatic examples that are not representative. Greater realism also makes ground truth harder to know. Reports should describe every intervention, compare clean controls, separate capability from propensity, test multiple model families when possible, and avoid using an existence proof as a frequency estimate or incident claim.

Sources: s1, s2, s3

Related terms

References

Last updated: 2026-09-04

In the Skills Atlas

This term is also covered in the Skills Atlas as model evaluation skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as adversarial ai testing skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as ai risk management skill.