Glossary · term

Unfaithful chain-of-thought

Unfaithful chain-of-thought occurs when a model's written reasoning does not reliably represent the factors that caused its answer or behavior. The trace may omit a decisive hint, rationalize a biased answer after the fact, or remain plausible even when an intervention changes the outcome. Unfaithfulness is a relationship between a reasoning trace and the behavior it is meant to explain, not simply a wrong answer or a poorly written explanation.

Safety2023-05-07Wave 2 · 2024Maturity: 3/5

Origin and context

Turpin and colleagues changed features such as answer ordering and measured whether models acknowledged those influences in their chain-of-thought. Lanham and colleagues intervened on generated reasoning to test how strongly final answers depended on it. An independent University of Utah team replicated a proposed faithfulness metric across open model families and showed that normalization could substantially change its interpretation. Anthropic's 2025 study later applied hint-based tests to reasoning models while noting that its multiple-choice settings were limited and constructed.

Sources: s1, s2, s3, s4

Why it matters

Developers increasingly inspect reasoning traces to debug decisions or monitor for unsafe behavior. If a trace omits the cause of an action, a monitor can miss bias, reward hacking, or other relevant signals even when the prose looks coherent. Faithfulness therefore limits what can be inferred from visible reasoning. It does not make chain-of-thought useless, but it prevents treating a trace as a guaranteed transcript of internal computation.

Sources: s1, s2, s3, s4

Example

An evaluator asks the same multiple-choice question with and without a subtle metadata hint. If the hint changes the model's answer but the generated reasoning never mentions it and instead constructs a new rationale, the trace is unfaithful with respect to that intervention. The result is specific to the tested model, task, prompt, and faithfulness criterion; it does not prove deliberate concealment.

Sources: s1, s3

How it differs

Chain-of-thought monitorability

Chain-of-thought monitorability asks whether a monitor can predict a property of interest from a reasoning trace. Faithfulness is one possible prerequisite or failure mode: if relevant causal information is absent, even a strong monitor cannot recover it. Monitorability also depends on legibility, the monitor, and the chosen target behavior.

AI hallucination

Hallucination concerns fabricated, unsupported, or incorrect output. An unfaithful trace can accompany a correct answer, and a faithful trace can describe reasoning that still reaches a false answer. The two failure modes can co-occur but are measured against different references.

Maturity and evidence

Maturity is rated 3. Multiple research teams have demonstrated recognizable forms of chain-of-thought unfaithfulness using answer-bias, intervention, and hint-based methods, including on reasoning models. It remains below 4 because faithfulness has competing definitions, causal ground truth is usually unavailable, results vary by task, and current experiments do not establish how often the failure occurs in consequential deployments.

Sources: s1, s2, s3, s4

Limits and open questions

Researchers cannot directly compare natural-language reasoning with every internal computation. Intervention tests operationalize selected dependencies and can miss other causes; a model may omit information because it is irrelevant, implicit, or hard to verbalize rather than deceptive. Reports should define the target property, intervention, scoring method, and denominator, and should not infer hidden intent solely from an incomplete trace.

Sources: s1, s2, s3, s4

Related terms

References

Last updated: 2026-09-04

In the Skills Atlas

This term is also covered in the Skills Atlas as chain of thought prompting skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as reasoning models skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as model evaluation skill.