Atlas · GenAI 2026

Explainable AI

Explainable AI (XAI) for GenAI

conceptPeak: 2025Explainability & FairnessAI consensus: 1/3

Prerequisites

  • Explaining Transformer decisions (attention visualization, probing, feature attribution) requires understanding the architecture

Recommended reference

Bills et al. (2023) 'Language Models Can Explain Neurons in Language Models' — OpenAI; mechanistic interpretability for Transformers

Notes from AI deep research

Anthropic Opus

Bills (2023) 'LMs Can Explain Neurons'. Mechanistic interpretability. W EU: prawo do wyjasnienia

Google Deep Think

Mechanistyczna interpretowalność [G#87]

Related skills