Summary (Overview)
- Core question: Does MUON, a matrix-momentum optimizer for language models, need fine-grained spectral shaping of its update direction, or is a coarse two-band reweighting sufficient?
- Key finding: A single shared bulk-to-spike gain (BULKBOOST) captures at least as much benefit as fine-grained spectral profiles like FREON and SPECTRA, suggesting useful departures from MUON's flat spectrum are surprisingly low-dimensional.
- Spectral diagnostics: Approximately 93.7–97.1% of singular modes of MUON's momentum lie below a Marchenko–Pastur noise edge; these bulk directions are weak individually but collectively aligned with a reference gradient.
- Empirical results: Across 30 continued-pretraining settings (Pythia-14M to 410M, six corpora), two-band reweighting reduces final loss by 0.073–0.147% of pre-adaptation loss vs. MUON, outperforming FREON's 0.022% improvement.
- Theoretical contribution: Provides a first-order condition for when shifting weight toward the bulk lowers loss, and quantifies the fraction of maximal improvement rate that two bands can capture.
Introduction and Theoretical Foundation
MUON combines current and past gradients into a matrix momentum . Its SVD decomposes the momentum into orthogonal directions weighted by singular values. The idealized polar update gives every singular direction unit gain—the flat profile.
Recent optimizers depart from this flat profile using spectral maps where gains vary across singular directions (FREON, SPECTRA, DEVA, and others). However, Shumaylov et al. (2026) showed that substantially different spectra, including random and inverted ones, can remain effective, raising the question: does MUON need fine-grained spectral shaping?
Two key observations motivate the two-band approach:
-
Signal vs. strength are distinct: A direction's singular value strength, its separation from sampling noise, and the benefit of increasing its update weight are different quantities. Small singular values can carry signal; large ones need not deserve more weight.
-
Spectral bulk structure: Under a Gaussian white-noise reference model, random matrix theory (Marchenko & Pastur, 1967; Baik et al., 2005) provides an asymptotic upper spectral edge. On Pythia-70M MUON trajectories, 93.7–97.1% of pooled singular modes lie in the bulk (below this edge), yet these weak directions collectively show positive signed alignment with an independently sampled reference gradient.
Methodology
MUON Setup and Notation
For a matrix parameter with minibatch gradient , the EMA buffer and Nesterov spectral input are:
with and .
Two-Band Partitions
BULKBOOST partitions singular directions into spike band and bulk band :
- BB-STATIC: Fixed-rank rule with for a fixed fraction (shared configuration uses )
- BB-ADAPTIVE: Noise-calibrated rule comparing singular values against threshold :
Norm-Matched Spectral Reweighting
The exact norm-matched direction for profile is:
BULKBOOST's two-level profile with bulk-to-spike gain :
Noise-Edge Calibration (BB-ADAPTIVE)
Using two independent minibatch halves, the split-difference estimates noise:
The threshold uses the propagated noise variance in the Nesterov input:
SVD-Free Realization
An exact polar-plus-low-rank identity enables avoiding full SVD:
where and . Spike-subspace tracking uses the smaller Gram side with warm-started block power iterations.
Empirical Validation / Results
Theoretical Results
Theorem 1 (Local spectral improvement): For a fixed partition with spike and bulk directions, with band means and and gap :
(i) Increasing from 1 gives strictly positive first-order improvement iff .
(ii) Two bands reach fraction of the best per-direction improvement rate.
Theorem 2 (Exact polar-plus-low-rank decomposition): The unnormalized spectral update satisfies:
Experimental Results
Setup: Continued pretraining of Pythia checkpoints (14M–410M) on six corpora (FineMath-4, FineMath-3, CodeParrot, Wiki-de, Wiki-zh, Wiki-ar), ~100M tokens per run, three seeds.
Key results from Table 1 (final validation loss, mean ± std over three seeds):
| Model | Method | FineMath-4 | Wiki-zh | Wiki-ar |
|---|---|---|---|---|
| 14M | MUON | 3.086 ± 0.002 | 2.745 ± 0.001 | 2.006 ± 0.000 |
| 14M | FREON | 3.083 ± 0.002 | 2.742 ± 0.001 | 2.008 ± 0.001 |
| 14M | BB-ADAPTIVE | 3.085 ± 0.002 | 2.737 ± 0.001 | 2.000 ± 0.001 |
| 70M | MUON | 2.551 ± 0.002 | 2.211 ± 0.005 | 1.702 ± 0.000 |
| 70M | FREON | 2.536 ± 0.001 | 2.200 ± 0.000 | 1.694 ± 0.001 |
| 70M | BB-ADAPTIVE | 2.545 ± 0.001 | 2.205 ± 0.004 | 1.690 ± 0.000 |
| 410M | MUON | 2.046 ± 0.005 | 1.879 ± 0.002 | 1.452 ± 0.001 |
| 410M | FREON | 2.043 ± 0.005 | 1.883 ± 0.001 | 1.460 ± 0.001 |
| 410M | BB-ADAPTIVE | 2.042 ± 0.005 | 1.875 ± 0.001 | 1.453 ± 0.002 |
Summary statistics (Table 2):
| Method | Wins vs. MUON | Mean Δ vs. MUON (% of ℓ₀) | Tokens saved (%) |
|---|---|---|---|
| FREON | 19/30 | -0.022 ± 0.010 | 7.78 |
| BB-STATIC(5%) | 23/30 | -0.073 ± 0.010 | 7.81 |
| BB-ADAPTIVE | 24/30 | -0.077 ± 0.010 | 7.99 |
| BB-STATIC† | 25/30 | -0.147 ± 0.021 | 10.00 |
Key findings:
- BB-STATIC(5%) beats FREON in 13 settings, loses in 13; BB-ADAPTIVE beats FREON in 17, loses in 10
- SPECTRA and DEVA trail MUON in all 30 settings
- BB-STATIC† beats BB-ADAPTIVE in 20 settings but requires loss-based fraction selection
- BB-ADAPTIVE matches BB-STATIC(5%) without fraction search (10 wins, 6 losses, 14 ties)
- Per-step runtimes: 1.36–1.50× MUON for BB-ADAPTIVE, 1.19–1.27× for static configurations
Theoretical and Practical Implications
Theoretical implications:
- Near the flat profile, any norm-matched reweighting moves along some direction in the tangent space , with best rate
- Two bands can only use the difference between band means; finer shaping adds at most the share from within-band utility differences
- The polar-plus-low-rank identity (Theorem 2) shows the correction for two-band reweighting is inherently low-rank, enabling efficient SVD-free implementation
Practical implications:
- A single shared gain parameter suffices for competitive performance, eliminating the need for per-direction gain tuning
- BB-ADAPTIVE provides automatic noise-calibrated splitting without hyperparameter search, at moderate computational cost
- Token savings of 8–10% to reach MUON's final loss suggest meaningful efficiency gains in continued pretraining
- The approach transfers across model scales (14M–410M) and diverse corpora (math, code, multilingual text)
Conclusion
The paper demonstrates that useful departures from MUON's flat spectral profile are surprisingly low-dimensional. A single bulk-to-spike gain captures at least as much benefit as fine-grained spectral profiles across 30 continued-pretraining settings. The key insights are:
-
Spectral structure matters but coarsely: The bulk (94–97% of modes below the noise edge) is weak individually but collectively useful, justifying a two-band treatment.
-
Theory guides practice: Theorem 1 provides a local test for when bulk emphasis helps; Theorem 2 enables efficient implementation without full SVD.
-
Competitive performance: Two-band reweighting matches or exceeds fine-grained baselines in aggregate, with BB-ADAPTIVE providing automatic split selection.
Future directions:
- Extending evaluation to other architectures and training regimes beyond Pythia continued pretraining
- Improving the noise-edge reference model beyond white Gaussian assumptions
- Exploring adaptive gain selection to replace the fixed default
- Investigating whether the two-band insight transfers to other optimizer families
Limitations acknowledged: The strongest BB-STATIC results require per-setting fraction selection; the noise edge only approximately separates signal from noise; evaluation is restricted to Pythia continued pretraining under the reported tuning protocol.
Related papers
- hacktrace: behavior-supervised detection of reward hacking during code generation
HACKTRACE detects reward hacking in coding agents from activations already computed during generation, cutting cheating from 85% to under 5% with only 8 ms overhead.
- Stateless Language Agents: Scaling Long-Horizon Automated Research
Stateless Language Agents, where the harness owns all research state and reconstructs fresh contexts per invocation, outperform stateful agent frameworks on long-horizon tasks, reaching baseline final performance with over 84% fewer tokens.
- VFold: Symmetry-Aware Cross-Layer Value Cache Compression
VFOLD compresses LLM value caches by folding cross-layer alignment maps into attention weights, achieving 25% KV reduction with over 98% performance retention.