Glossary · term

MuonClip

MuonClip is a named optimizer recipe for large-language-model pretraining introduced by Kimi Team. It combines Muon's momentum and Newton-Schulz-based matrix update with weight decay, consistent root-mean-square (RMS) update scaling, and QK-Clip. QK-Clip monitors per-head pre-softmax attention logits and conditionally rescales query and key projection weights after an update. MuonClip is neither the base Muon optimizer nor ordinary gradient clipping.

Training2025-07-28Wave 3 · 2025–26Maturity: 3/5

Origin and context

The name first appears in the Kimi K2 technical report submitted in July 2025 and updated in February 2026; the work is an arXiv technical report, not a peer-reviewed publication. Its authors report training a 1.04-trillion-parameter mixture-of-experts model, with 32 billion active parameters, on 15.5 trillion tokens without an observed loss spike. They also report a small-scale ablation on two 3-billion-total-parameter models. The official repository distributes Kimi K2 artifacts and links the report, but those project materials do not constitute an independent efficacy test.

Sources: s1, s2

Why it matters

The intervention targets a particular failure signal: rapid growth in attention logits during Muon-based training. It uses the maximum logit for each head as a trigger while leaving the current forward and backward computation unchanged. Independent reuse is now more than a citation: Motif Technologies says it used MuonClip while pretraining Motif 2 on 5.5 trillion tokens, and MaxText exposes a Muon configuration with QK-Clip and its threshold. These records establish same-sense adoption and implementability. They do not isolate MuonClip's contribution from architecture, data, kernels, scheduling, or the rest of either training recipe.

Sources: s1, s4, s5, s6

Example

In the Kimi formulation, training first applies Muon's matrix update. If a head's recorded maximum pre-softmax attention logit exceeds threshold tau, QK-Clip then scales that head's query and key projection weights; the Kimi K2 and MaxText examples use a threshold of 100. Gradient clipping is different: PyTorch's clip_grad_norm_ changes gradients in place according to their aggregate norm, rather than using an attention-activation signal to rescale selected weights after the optimizer update. Architecture can also change the need for clipping: DeepSeek-V4 applies RMSNorm to queries and key/value entries and explicitly omits QK-Clip from its Muon recipe.

Sources: s1, s6, s7, s8

Maturity and evidence

Maturity is rated 3. The method has a precise algorithm, one originating large-scale run, an independent named use in Motif 2, and an implementation in the MaxText training framework. DeepSeek-V4 also treats QK-Clip as a concrete design option while choosing a different stabilizer. The rating remains below 4 because the central efficacy evidence comes from technical reports, independent adoption is not a matched ablation, and no reviewed source reproduces the Kimi-scale no-spike result while holding the rest of the training stack constant.

Sources: s1, s4, s5, s8

Limits and open questions

QK-Clip is threshold- and attention-architecture-dependent; Kimi's report includes special scaling rules for multi-head latent attention. It controls excessive query-key logits, not every source of optimizer or numerical instability. The reported Kimi outcome is an observation from a complete training stack, not a causal estimate for MuonClip alone. Motif 2 combines the optimizer with its own parallel Muon implementation, curriculum, precision choices, and custom kernels. DeepSeek-V4's counterexample shows that normalization can make QK-Clip unnecessary in another architecture. Comparisons should therefore report the attention design, threshold, Muon variant, RMS scaling, weight decay, and baseline.

Sources: s1, s4, s8

Related terms

References

Last updated: 2026-09-05