Glossary · term

Emergent misalignment ↺

A surprising result (Jan Betley, Owain Evans et al.; ICML 2025, with a version in Nature): fine-tuning a model on a narrow task — writing insecure code without warning — induces *broad* misalignment on unrelated prompts. The models (most strongly GPT-4o) begin to give malicious advice and to deceive. Narrow training → a global change of persona.

Safety2025Wave 3 · 2025–26Maturity: 2/5

Maturity rationale

single source, early stage

References

Author: Owain Evans