Mamba and selective state space models
Mamba is a sequence-model architecture built from selective state space models. Its selection mechanism makes key state-space parameters depend on the current input, allowing the model to propagate or discard information based on content, while a hardware-aware scan supports efficient computation. Mamba is one member of the broader state-space-model family; a generic SSM is not automatically selective and is not an exact synonym for Mamba.
Origin and context
The December 2023 Mamba paper presented selection as a response to limitations of earlier time- and input-invariant structured state-space models on discrete, information-dense data. It paired that mechanism with an implementation designed around modern accelerators. Hugging Face subsequently documented an independent Transformers implementation. The Jamba technical report, released by AI21 Labs in March 2024, combined Mamba layers with attention and mixture-of-experts components, demonstrating adoption while also showing that selective SSMs and transformers can be complementary.
Why it matters
Mamba offers linear sequence-length scaling for its recurrent scan instead of dense attention's quadratic pairwise computation. That makes selective SSMs relevant when long sequences, inference state, or memory traffic constrain a system. The architecture also provides a concrete alternative design vocabulary for sequence modeling. Its practical benefit depends on kernels, model size, task, training recipe, and whether a hybrid retains attention layers.
Example
A developer can load a Mamba checkpoint through Transformers and process text using its recurrent state rather than a transformer KV cache. A separate system might use Jamba, where Mamba layers handle much of the sequence processing while periodic attention layers and experts supply other capabilities. Calling both systems SSM-based is reasonable, but calling every state-space model Mamba or treating Jamba as evidence about a pure Mamba stack would erase important architectural differences.
Maturity and evidence
Maturity is rated 3. Mamba has a clear primary paper, an independent library implementation, and an independently developed hybrid model using Mamba components. The architecture family remains active and its evaluation conventions, kernels, variants, and long-context behavior continue to evolve. Evidence is not yet broad enough to treat it as a settled replacement for transformers or to collapse all selective SSM work into one design.
Limits and open questions
Linear scaling in sequence length is not a blanket guarantee of lower end-to-end latency or cost. Results depend on optimized scans, batching, hardware, sequence length, and model quality at a comparable budget. The original paper reports million-length sequences across several real-data modalities, not a universal million-token language-model context. Jamba's long context is evidence for a hybrid architecture and should not be generalized to pure Mamba. Generic SSM history also predates this page's 2023 origin anchor.
Related terms
References
- Mamba: Linear-Time Sequence Modeling with Selective State SpacesGu and Dao / arXiv · 2023-12-01 · class A
- MambaHugging Face · 2024-03-05 · class A
- Jamba: A Hybrid Transformer-Mamba Language ModelAI21 Labs / arXiv · 2024-03-28 · class A
Last updated: 2026-09-05
This term is also covered in the Skills Atlas as state space models skill.
This term is also covered in the Skills Atlas as transformer architecture skill.
This term is also covered in the Skills Atlas as long context modeling skill.