Glossary · term

Chain-of-thought monitorability ↺

The thesis that reasoning models that “think” in natural language offer a rare opportunity for safety oversight: their chain of thought can be monitored for intent to do harm. The authors stress that this window is fragile and that training decisions may inadvertently close it, and so they call for protecting it.

Safety2025Wave 3 · 2025–26Maturity: 2/5

Maturity rationale

single source, early stage

References

Author: Geoffrey Hinton