Glossary · term
Chain-of-thought monitorability
A safety doctrine holding that the chain of thought (CoT) of frontier models can be monitored for signs of intent to misbehave, providing a valuable but fragile layer of oversight. It works only as long as the CoT remains legible; opaque RL and CoT compression can destroy it. An issue brief by more than 40 researchers (2025).
SafetyI 2026Wave 2 · 2024Maturity: 1/5
Maturity rationale
Agent runaway — a neologism, a scenario not a standard
References
Author: Frontier Model Forum