Atlas · GenAI 2026
Mechanistic Interpretability
Reverse-engineering model internals into human-understandable circuits and features (sparse autoencoders, activation patching).
conceptPeak: 2024Explainability & FairnessAI consensus: 0/3
Prerequisites
- hardDeep Learning
You inspect the weights and activations of a network.
- mediumLinear Algebra
Features and circuits are analyzed in activation space.
Recommended reference
Towards Monosemanticity: Decomposing Language Models With Dictionary Learning — Anthropic, 2023