Atlas · GenAI 2026

Mixture of Experts

Mixture of Experts (MoE)

conceptPeak: 2024Neural ArchitecturesAI consensus: 1/3

Prerequisites

  • MoE replaces the dense MLP block in a Transformer with routed sparse experts — you must understand the Transformer to modify it

Recommended reference

Jiang et al. (2024) 'Mixtral of Experts' — arXiv:2401.04088; practical MoE at scale; plus Fedus et al. (2022) 'Switch Transformers' for theory

Notes from AI deep research

Anthropic Opus

Mixtral pokazal ze MoE dziala. Sparse routing = mniej compute przy zachowaniu capacity

Google Deep Think

Modele rzadkie; routing tokenów [G#23]

Related skills