Atlas · GenAI 2026

Mechanistic Interpretability

Reverse-engineering model internals into human-understandable circuits and features (sparse autoencoders, activation patching).

conceptPeak: 2024Explainability & FairnessAI consensus: 0/3

Prerequisites

  • You inspect the weights and activations of a network.

  • Features and circuits are analyzed in activation space.

Recommended reference

Towards Monosemanticity: Decomposing Language Models With Dictionary Learning — Anthropic, 2023

Notes from AI deep research