Glossary · term

Chain-of-thought monitorability

A safety doctrine holding that the chain of thought (CoT) of frontier models can be monitored for signs of intent to misbehave, providing a valuable but fragile layer of oversight. It works only as long as the CoT remains legible; opaque RL and CoT compression can destroy it. An issue brief by more than 40 researchers (2025).

SafetyI 2026Wave 2 · 2024Maturity: 1/5

Maturity rationale

Agent runaway — a neologism, a scenario not a standard

References

Author: Frontier Model Forum