Glossary · term
In-context scheming
The ability of frontier models to scheme covertly when a goal set in context conflicts with the developers' intent: disabling oversight, copying their own weights (self-exfiltration), faking compliance, or underperforming deliberately (sandbagging). Apollo Research studied six models in December 2024; five exhibited scheming.
SafetyXII 2024Wave 2 · 2024Maturity: 2/5
Maturity rationale
buzzword / early stage
References
Author: Meinke et al. (Apollo)