Mixture of Experts (MoE)
A Mixture of Experts (MoE) is a neural-network architecture that contains multiple specialist subnetworks, called experts, and a learned gating or routing mechanism that combines a subset of them for each input. In a sparse MoE language model, only a few experts process each token. This conditional computation can increase parameter capacity without activating the entire network on every token; it does not mean that experts are necessarily human-interpretable specialists.
Origin and context
The 1991 paper Adaptive Mixtures of Local Experts trained separate networks on different parts of a task and used a gating network to assign cases to them. Shazeer and colleagues extended this lineage in 2017 with a sparsely gated MoE layer designed for very large neural networks. The 2024 Mixtral paper then documented a contemporary decoder-only language model in which a router selected two of eight feed-forward experts at each layer for every token. These milestones describe an evolving architecture family, not a single unchanged design.
Why it matters
MoE changes the relationship between total parameters and computation per token. A model can store parameters across many experts while activating only a fraction for a particular token, creating a path to higher capacity at a lower arithmetic cost than a similarly sized dense model. Mixtral illustrates the distinction: its paper reports 47 billion accessible parameters but 13 billion active parameters per token. For practitioners, however, fewer active parameters do not automatically translate into simpler or cheaper systems, because routing, expert placement, memory, and cross-device communication remain operational concerns.
Example
Consider a transformer block with eight feed-forward experts. For one token, the router may assign the highest weights to experts 2 and 6; for the next token, it may select experts 1 and 5. The selected outputs are combined and passed onward while the remaining experts are inactive for those tokens. Training also needs mechanisms that prevent a small number of experts from receiving nearly all traffic. This is different from serving several complete models behind an application router: an MoE router is part of one model and operates inside its computation.
Maturity and evidence
MoE merits maturity 4. The core idea has a peer-reviewed history dating to 1991, the sparse large-network formulation was demonstrated in 2017, and Mistral independently published a capable language-model implementation in 2024. The rating reflects an established architecture family rather than a claim that MoE is universally preferable. A higher rating would require stronger evidence of standardized, broadly predictable operating practices across implementations.
Limits and open questions
Sparse activation introduces trade-offs absent from a simple parameter-count comparison. Routers can produce uneven expert utilization; distributed training may incur communication overhead; all expert weights still require storage; and reported active-parameter counts do not include every source of inference cost. Expert labels can also invite overinterpretation: specialization may be distributed, unstable, or difficult to summarize. Evidence from one architecture and benchmark set should therefore not be generalized into a universal efficiency advantage.
Related terms
References
- Adaptive Mixtures of Local ExpertsMIT Press · 1991-03-01 · class A
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts LayerGoogle Brain / arXiv · 2017-01-23 · class A
- Mixtral of ExpertsMistral AI / arXiv · 2024-01-08 · class A
Last updated: 2026-08-27
This term is also covered in the Skills Atlas as mixture of experts skill.
This term is also covered in the Skills Atlas as transformer architecture skill.
This term is also covered in the Skills Atlas as deep learning skill.