Atlas · GenAI 2026
Mixture of Experts
Mixture of Experts (MoE)
conceptPeak: 2024Neural ArchitecturesAI consensus: 1/3
Prerequisites
MoE replaces the dense MLP block in a Transformer with routed sparse experts — you must understand the Transformer to modify it
Recommended reference
Jiang et al. (2024) 'Mixtral of Experts' — arXiv:2401.04088; practical MoE at scale; plus Fedus et al. (2022) 'Switch Transformers' for theory
Notes from AI deep research
Anthropic Opus
Mixtral pokazal ze MoE dziala. Sparse routing = mniej compute przy zachowaniu capacity
Google Deep Think
Modele rzadkie; routing tokenów [G#23]
Related skills
- → is subcategory of: Deep Learning(3/3)
- → is an instance of: Transformer Architecture(1/3)