MuonClip
MuonClip is a named optimizer recipe for large-language-model pretraining introduced by Kimi Team. It combines Muon's momentum and Newton-Schulz-based matrix update with weight decay, consistent root-mean-square (RMS) update scaling, and QK-Clip. QK-Clip monitors per-head pre-softmax attention logits and conditionally rescales query and key projection weights after an update. MuonClip is neither the base Muon optimizer nor ordinary gradient clipping.
Origin and context
The name first appears in the Kimi K2 technical report submitted in July 2025 and updated in February 2026; the work is an arXiv technical report, not a peer-reviewed publication. Its authors report training a 1.04-trillion-parameter mixture-of-experts model, with 32 billion active parameters, on 15.5 trillion tokens without an observed loss spike. They also report a small-scale ablation on two 3-billion-total-parameter models. The official repository distributes Kimi K2 artifacts and links the report, but those project materials do not constitute an independent efficacy test.
Why it matters
The intervention targets a particular failure signal: rapid growth in attention logits during Muon-based training. It uses the maximum logit for each head as a trigger while leaving the current forward and backward computation unchanged. Independent reuse is now more than a citation: Motif Technologies says it used MuonClip while pretraining Motif 2 on 5.5 trillion tokens, and MaxText exposes a Muon configuration with QK-Clip and its threshold. These records establish same-sense adoption and implementability. They do not isolate MuonClip's contribution from architecture, data, kernels, scheduling, or the rest of either training recipe.
Example
In the Kimi formulation, training first applies Muon's matrix update. If a head's recorded maximum pre-softmax attention logit exceeds threshold tau, QK-Clip then scales that head's query and key projection weights; the Kimi K2 and MaxText examples use a threshold of 100. Gradient clipping is different: PyTorch's clip_grad_norm_ changes gradients in place according to their aggregate norm, rather than using an attention-activation signal to rescale selected weights after the optimizer update. Architecture can also change the need for clipping: DeepSeek-V4 applies RMSNorm to queries and key/value entries and explicitly omits QK-Clip from its Muon recipe.
Maturity and evidence
Maturity is rated 3. The method has a precise algorithm, one originating large-scale run, an independent named use in Motif 2, and an implementation in the MaxText training framework. DeepSeek-V4 also treats QK-Clip as a concrete design option while choosing a different stabilizer. The rating remains below 4 because the central efficacy evidence comes from technical reports, independent adoption is not a matched ablation, and no reviewed source reproduces the Kimi-scale no-spike result while holding the rest of the training stack constant.
Limits and open questions
QK-Clip is threshold- and attention-architecture-dependent; Kimi's report includes special scaling rules for multi-head latent attention. It controls excessive query-key logits, not every source of optimizer or numerical instability. The reported Kimi outcome is an observation from a complete training stack, not a causal estimate for MuonClip alone. Motif 2 combines the optimizer with its own parallel Muon implementation, curriculum, precision choices, and custom kernels. DeepSeek-V4's counterexample shows that normalization can make QK-Clip unnecessary in another architecture. Comparisons should therefore report the attention design, threshold, Muon variant, RMS scaling, weight decay, and baseline.
Related terms
References
- Kimi K2: Open Agentic IntelligenceKimi Team / arXiv · 2025-07-28 · class A
- Kimi K2Moonshot AI · 2025-07 · class A
- Muon is Scalable for LLM TrainingMoonshot AI and collaborators / arXiv · 2025-02-24 · class A
- Motif 2 12.7B technical reportMotif Technologies / arXiv · 2025-11-07 · class A
- MaxText release maxtext-v0.2.2Google Cloud AI Hypercomputer · 2026-05-08 · class A
- Run Kimi models with MaxTextGoogle Cloud AI Hypercomputer · 2026-05 · class A
- torch.nn.utils.clip_grad_norm_PyTorch · 2026 · class A
- DeepSeek-V4: Towards Highly Efficient Million-Token Context IntelligenceDeepSeek-AI / arXiv · 2026-04-26 · class A
Last updated: 2026-09-05