Glossary · term

Kimi Linear

Kimi Linear is a named hybrid language-model architecture introduced by the Kimi Team. Its reference configuration interleaves Kimi Delta Attention, or KDA, with Multi-Head Latent Attention layers. KDA is the recurrent linear-attention component and extends a gated delta-rule design with more fine-grained gating; it is not an alias for the whole architecture. The entry covers the reusable architecture and its implementation pattern, not the Kimi product family or every model that uses linear attention.

Training2025-10-30Wave 3 · 2025–26Maturity: 3/5

Origin and context

The originating Kimi Linear technical report was posted as a preprint on 30 October 2025. It defines the KDA module, the hybrid layer schedule and a 48-billion-parameter mixture-of-experts instantiation, then reports efficiency and quality results from the proposing team. NVIDIA's NeMo AutoModel 0.4 documentation exposes separate KimiDeltaAttention, KimiMLAAttention and KimiLinear model classes, which independently confirms that the named architecture can be implemented outside its origin repository. A 2026 independent preprint studies fast, stable triangular inversion for delta-rule linear transformers and includes Kimi Linear among the relevant open models.

Sources: s1, s2, s3

Why it matters

Kimi Linear is a concrete test of a broader design strategy: use recurrent or linear attention for most token processing while retaining selected softmax-attention layers for capabilities that benefit from direct token-to-token access. Its importance is architectural rather than product-based because the layer pattern, KDA recurrence and implementation interfaces can be studied and reimplemented independently. It does not establish that one hybrid ratio is universally optimal, and the proposing team's throughput, memory and benchmark comparisons should not be generalized across hardware, context lengths or training regimes.

Sources: s1, s2, s3

Example

In a Kimi Linear block schedule, most layers can use KDA to update a compact recurrent state, while occasional MLA layers retain explicit softmax attention. A framework implementation therefore needs distinct KDA state-update code, MLA attention code and the configuration that interleaves them. Calling KDA alone 'Kimi Linear' loses that composition: KDA is a reusable core module, whereas Kimi Linear names the full hybrid architecture and its prescribed family of model configurations.

Sources: s1, s3

How it differs

Hybrid Attention Architecture

Hybrid attention architecture is the broader pattern of mixing softmax-attention layers with recurrent, state-space or linear-attention token mixers. Kimi Linear is a specific architecture within that pattern, with its own KDA and MLA components; the broader category and the named architecture are not synonymous.

Mamba and selective state space models

Mamba is a selective state-space architecture. KDA instead uses a gated delta-rule linear-attention recurrence; both avoid full attention in some layers, but their state updates are not interchangeable.

Maturity and evidence

Maturity is rated 3. The architecture is specified in an originating preprint, implemented by an independent model framework and discussed in independent numerical-method research. This is stronger than a single model announcement. Lifecycle remains emerging because the name is young, independent evidence focuses on implementation and one kernel-level issue, and comparable-scale replication of the origin report's end-to-end results is not yet established.

Sources: s1, s2, s3

Limits and open questions

The headline quality and efficiency results come from the proposing team and depend on its exact model, kernels and hardware. Independent NeMo support demonstrates portability, not benchmark superiority. Delta-rule implementations can face numerical-stability and triangular-inversion trade-offs, which the independent preprint analyzes. The name still identifies a particular proposed architecture; initial framework support is not evidence of broad use or comparable end-to-end performance.

Sources: s1, s2, s3

Related terms

References

Last updated: 2026-09-05

In the Skills Atlas

This term is also covered in the Skills Atlas as transformer architecture skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as long context modeling skill.