Glossary · term

Circuit Tracing

Circuit tracing is an emerging mechanistic-interpretability method for constructing a prompt-specific graph of how internal features contribute to a model output. Anthropic's 2025 method replaces selected model components with cross-layer transcoders trained to approximate them, then computes an attribution graph over interpretable features. The graph is a model-assisted hypothesis about computation, not a complete trace of every operation in the original neural network.

Safety2025-03-27Wave 3 · 2025–26Maturity: 2/5

Origin and context

Anthropic introduced the named method in March 2025 in Circuit Tracing: Revealing Computational Graphs in Language Models. The work described a replacement model, attribution graphs, visualization tools, and interventions for validating hypotheses on an 18-layer language model, with a companion application to Claude 3.5 Haiku. In February 2026, an independent paper extended a circuit-tracing framework to vision-language models using transcoders, attribution graphs, attention-based analysis, feature steering, and circuit patching. That extension is evidence of adoption, but the terminology and implementations remain young.

Sources: s1, s2

Why it matters

Circuit tracing turns selected internal influences into a navigable graph, which can help researchers form and test hypotheses about multi-step behaviors that are hard to localize to one neuron or one layer. The Anthropic work includes interventions intended to test whether graph components matter causally, while the independent VLM study applies related tools to multimodal reasoning. Such graphs may support targeted investigation and comparison, but they do not automatically certify a behavior, expose every causal path, or establish that feature labels are correct.

Sources: s1, s2

Example

For a prompt that elicits a factual answer, a researcher can generate an attribution graph, group related feature nodes, and identify paths that appear to connect the subject tokens to the output. The researcher then intervenes on selected features or patches a circuit and checks whether the answer changes as predicted. In a vision-language setting, the same pattern can test whether visual features contribute to a reasoning result. A static diagram without intervention or approximation checks is evidence visualization, not a validated computational account.

Sources: s1, s2

Maturity and evidence

Circuit tracing remains maturity 2. It has a detailed primary method and an independent 2026 extension into vision-language models. The evidence is still concentrated in recent research papers, with limited replication and no shared evaluation standard. Multiple independent implementations, benchmarked faithfulness, and stable results across model families would strengthen the maturity assessment.

Sources: s1, s2

Limits and open questions

The replacement model only approximates the original computation, graph construction requires pruning and attribution choices, and feature labels may be incomplete or misleading. A prompt-specific graph need not generalize to paraphrases, tasks, or checkpoints. Interventions can also introduce behavior outside the model's usual activation distribution. An attribution graph is an output representation used in this method, not an independently validated explanation merely because it is visually interpretable.

Sources: s1, s2

Related terms

References

Last updated: 2026-09-05

In the Skills Atlas

This term is also covered in the Skills Atlas as mechanistic interpretability skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as transformer architecture skill.