Summary (Overview)

  • Core question: Does MUON, a matrix-momentum optimizer for language models, need fine-grained spectral shaping of its update direction, or is a coarse two-band reweighting sufficient?
  • Key finding: A single shared bulk-to-spike gain (BULKBOOST) captures at least as much benefit as fine-grained spectral profiles like FREON and SPECTRA, suggesting useful departures from MUON's flat spectrum are surprisingly low-dimensional.
  • Spectral diagnostics: Approximately 93.7–97.1% of singular modes of MUON's momentum lie below a Marchenko–Pastur noise edge; these bulk directions are weak individually but collectively aligned with a reference gradient.
  • Empirical results: Across 30 continued-pretraining settings (Pythia-14M to 410M, six corpora), two-band reweighting reduces final loss by 0.073–0.147% of pre-adaptation loss vs. MUON, outperforming FREON's 0.022% improvement.
  • Theoretical contribution: Provides a first-order condition for when shifting weight toward the bulk lowers loss, and quantifies the fraction of maximal improvement rate that two bands can capture.

Introduction and Theoretical Foundation

MUON combines current and past gradients into a matrix momentum MtM_t. Its SVD Mt=UtΣtVt⊤M_t = U_t \Sigma_t V_t^\top decomposes the momentum into orthogonal directions weighted by singular values. The idealized polar update Qt=polar(Mt)=UtVt⊤Q_t = \mathrm{polar}(M_t) = U_t V_t^\top gives every singular direction unit gain—the flat profile.

Recent optimizers depart from this flat profile using spectral maps Ug(Σ)V⊤U g(\Sigma) V^\top where gains vary across singular directions (FREON, SPECTRA, DEVA, and others). However, Shumaylov et al. (2026) showed that substantially different spectra, including random and inverted ones, can remain effective, raising the question: does MUON need fine-grained spectral shaping?

Two key observations motivate the two-band approach:

  1. Signal vs. strength are distinct: A direction's singular value strength, its separation from sampling noise, and the benefit of increasing its update weight are different quantities. Small singular values can carry signal; large ones need not deserve more weight.

  2. Spectral bulk structure: Under a Gaussian white-noise reference model, random matrix theory (Marchenko & Pastur, 1967; Baik et al., 2005) provides an asymptotic upper spectral edge. On Pythia-70M MUON trajectories, 93.7–97.1% of pooled singular modes lie in the bulk (below this edge), yet these weak directions collectively show positive signed alignment with an independently sampled reference gradient.

Methodology

MUON Setup and Notation

For a matrix parameter Wt∈Rm×nW_t \in \mathbb{R}^{m \times n} with minibatch gradient GtG_t, the EMA buffer and Nesterov spectral input are:

Bt=βBt−1+(1−β)Gt,Mt=βBt+(1−β)Gt(1)B_t = \beta B_{t-1} + (1-\beta) G_t, \quad M_t = \beta B_t + (1-\beta) G_t \tag{1}

with β∈[0,1)\beta \in [0,1) and B0=0B_0 = 0.

Two-Band Partitions

BULKBOOST partitions singular directions into spike band Isp,t\mathcal{I}_{\mathrm{sp},t} and bulk band Ibu,t\mathcal{I}_{\mathrm{bu},t}:

  • BB-STATIC: Fixed-rank rule with rt=⌈pd⌉r_t = \lceil p d \rceil for a fixed fraction pp (shared configuration uses p=5%p = 5\%)
  • BB-ADAPTIVE: Noise-calibrated rule comparing singular values against threshold τt\tau_t:
Isp,t={i∈[d]:σt,i>τt},Ibu,t={i∈[d]:σt,i≤τt}(2)\mathcal{I}_{\mathrm{sp},t} = \{i \in [d]: \sigma_{t,i} > \tau_t\}, \qquad \mathcal{I}_{\mathrm{bu},t} = \{i \in [d]: \sigma_{t,i} \leq \tau_t\} \tag{2}

Norm-Matched Spectral Reweighting

The exact norm-matched direction for profile wtw_t is:

Dt(wt)=ct(wt)Utdiag⁡(wt)Vt⊤,ct(wt)=d∑i=1dwt,i2(3)D_t(w_t) = c_t(w_t) U_t \operatorname{diag}(w_t) V_t^\top, \qquad c_t(w_t) = \sqrt{\frac{d}{\sum_{i=1}^{d} w_{t,i}^2}} \tag{3}

BULKBOOST's two-level profile with bulk-to-spike gain ρ\rho:

wt,i(ρ)={1,i∈Isp,t,ρ,i∈Ibu,t,cρ,rt:=ct(wt(ρ))=drt+ρ2(d−rt)(4)w_{t,i}^{(\rho)} = \begin{cases} 1, & i \in \mathcal{I}_{\mathrm{sp},t}, \\ \rho, & i \in \mathcal{I}_{\mathrm{bu},t}, \end{cases} \quad c_{\rho,r_t} := c_t(w_t^{(\rho)}) = \sqrt{\frac{d}{r_t + \rho^2(d - r_t)}} \tag{4}

Noise-Edge Calibration (BB-ADAPTIVE)

Using two independent minibatch halves, the split-difference estimates noise:

s^t2=∥Ztsplit∥F2mn,sˉt2=γsˉt−12+(1−γ)χt2s^t2(6)\widehat{s}_t^2 = \frac{\|Z_t^{\mathrm{split}}\|_F^2}{mn}, \qquad \bar{s}_t^2 = \gamma \bar{s}_{t-1}^2 + (1-\gamma) \chi_t^2 \widehat{s}_t^2 \tag{6}

The threshold uses the propagated noise variance in the Nesterov input:

νt2(β)=(1−β)2[(1+β)2+β4(1−β2(t−1))1−β2](7)\nu_t^2(\beta) = (1-\beta)^2 \left[ (1+\beta)^2 + \frac{\beta^4(1-\beta^{2(t-1)})}{1-\beta^2} \right] \tag{7} τt=τ^t=κedgeνt(β)sˉt2(m+n)(8)\tau_t = \hat{\tau}_t = \kappa_{\mathrm{edge}} \nu_t(\beta) \sqrt{\bar{s}_t^2}(\sqrt{m} + \sqrt{n}) \tag{8}

SVD-Free Realization

An exact polar-plus-low-rank identity enables avoiding full SVD:

Dtsvd={cρ,rt[ρQt−(ρ−1)Pt,spQt],m≤n,cρ,rt[ρQt−(ρ−1)QtRt,sp],n<m.(9)D_t^{\mathrm{svd}} = \begin{cases} c_{\rho,r_t}[\rho Q_t - (\rho-1) P_{t,\mathrm{sp}} Q_t], & m \leq n, \\ c_{\rho,r_t}[\rho Q_t - (\rho-1) Q_t R_{t,\mathrm{sp}}], & n < m. \end{cases} \tag{9}

where Pt,sp=Ut,spUt,sp⊤P_{t,\mathrm{sp}} = U_{t,\mathrm{sp}} U_{t,\mathrm{sp}}^\top and Rt,sp=Vt,spVt,sp⊤R_{t,\mathrm{sp}} = V_{t,\mathrm{sp}} V_{t,\mathrm{sp}}^\top. Spike-subspace tracking uses the smaller Gram side with warm-started block power iterations.

Empirical Validation / Results

Theoretical Results

Theorem 1 (Local spectral improvement): For a fixed partition with rr spike and k=d−rk = d-r bulk directions, with band means uˉsp\bar{u}_{\mathrm{sp}} and uˉbu\bar{u}_{\mathrm{bu}} and gap Δ=uˉbu−uˉsp\Delta = \bar{u}_{\mathrm{bu}} - \bar{u}_{\mathrm{sp}}:

(i) Increasing ρ\rho from 1 gives strictly positive first-order improvement iff Δ>0\Delta > 0.

(ii) Two bands reach fraction Rbt=rk/d∣Δ∣∥u⊥∥2∈[0,1]R_{\mathrm{bt}} = \frac{\sqrt{rk/d}|\Delta|}{\|u_\perp\|_2} \in [0,1] of the best per-direction improvement rate.

Theorem 2 (Exact polar-plus-low-rank decomposition): The unnormalized spectral update satisfies:

D~(a,ρ)=ρQ+Uspdiag⁡(a−ρ1r)Vsp⊤,rank⁡(D~(a,ρ)−ρQ)≤r(14)\widetilde{D}(a,\rho) = \rho Q + U_{\mathrm{sp}} \operatorname{diag}(a - \rho \mathbf{1}_r) V_{\mathrm{sp}}^\top, \quad \operatorname{rank}(\widetilde{D}(a,\rho) - \rho Q) \leq r \tag{14}

Experimental Results

Setup: Continued pretraining of Pythia checkpoints (14M–410M) on six corpora (FineMath-4, FineMath-3, CodeParrot, Wiki-de, Wiki-zh, Wiki-ar), ~100M tokens per run, three seeds.

Key results from Table 1 (final validation loss, mean ± std over three seeds):

ModelMethodFineMath-4Wiki-zhWiki-ar
14MMUON3.086 ± 0.0022.745 ± 0.0012.006 ± 0.000
14MFREON3.083 ± 0.0022.742 ± 0.0012.008 ± 0.001
14MBB-ADAPTIVE3.085 ± 0.0022.737 ± 0.0012.000 ± 0.001
70MMUON2.551 ± 0.0022.211 ± 0.0051.702 ± 0.000
70MFREON2.536 ± 0.0012.200 ± 0.0001.694 ± 0.001
70MBB-ADAPTIVE2.545 ± 0.0012.205 ± 0.0041.690 ± 0.000
410MMUON2.046 ± 0.0051.879 ± 0.0021.452 ± 0.001
410MFREON2.043 ± 0.0051.883 ± 0.0011.460 ± 0.001
410MBB-ADAPTIVE2.042 ± 0.0051.875 ± 0.0011.453 ± 0.002

Summary statistics (Table 2):

MethodWins vs. MUONMean Δ vs. MUON (% of ℓ₀)Tokens saved (%)
FREON19/30-0.022 ± 0.0107.78
BB-STATIC(5%)23/30-0.073 ± 0.0107.81
BB-ADAPTIVE24/30-0.077 ± 0.0107.99
BB-STATIC†25/30-0.147 ± 0.02110.00

Key findings:

  • BB-STATIC(5%) beats FREON in 13 settings, loses in 13; BB-ADAPTIVE beats FREON in 17, loses in 10
  • SPECTRA and DEVA trail MUON in all 30 settings
  • BB-STATIC† beats BB-ADAPTIVE in 20 settings but requires loss-based fraction selection
  • BB-ADAPTIVE matches BB-STATIC(5%) without fraction search (10 wins, 6 losses, 14 ties)
  • Per-step runtimes: 1.36–1.50× MUON for BB-ADAPTIVE, 1.19–1.27× for static configurations

Theoretical and Practical Implications

Theoretical implications:

  • Near the flat profile, any norm-matched reweighting moves along some direction in the tangent space T={δ∈Rd:1d⊤δ=0}\mathcal{T} = \{\delta \in \mathbb{R}^d: \mathbf{1}_d^\top \delta = 0\}, with best rate ∥u⊥∥2\|u_\perp\|_2
  • Two bands can only use the difference between band means; finer shaping adds at most the share 1−Rbt1 - R_{\mathrm{bt}} from within-band utility differences
  • The polar-plus-low-rank identity (Theorem 2) shows the correction for two-band reweighting is inherently low-rank, enabling efficient SVD-free implementation

Practical implications:

  • A single shared gain parameter ρ\rho suffices for competitive performance, eliminating the need for per-direction gain tuning
  • BB-ADAPTIVE provides automatic noise-calibrated splitting without hyperparameter search, at moderate computational cost
  • Token savings of 8–10% to reach MUON's final loss suggest meaningful efficiency gains in continued pretraining
  • The approach transfers across model scales (14M–410M) and diverse corpora (math, code, multilingual text)

Conclusion

The paper demonstrates that useful departures from MUON's flat spectral profile are surprisingly low-dimensional. A single bulk-to-spike gain captures at least as much benefit as fine-grained spectral profiles across 30 continued-pretraining settings. The key insights are:

  1. Spectral structure matters but coarsely: The bulk (94–97% of modes below the noise edge) is weak individually but collectively useful, justifying a two-band treatment.

  2. Theory guides practice: Theorem 1 provides a local test for when bulk emphasis helps; Theorem 2 enables efficient implementation without full SVD.

  3. Competitive performance: Two-band reweighting matches or exceeds fine-grained baselines in aggregate, with BB-ADAPTIVE providing automatic split selection.

Future directions:

  • Extending evaluation to other architectures and training regimes beyond Pythia continued pretraining
  • Improving the noise-edge reference model beyond white Gaussian assumptions
  • Exploring adaptive gain selection to replace the fixed ρ=4\rho = 4 default
  • Investigating whether the two-band insight transfers to other optimizer families

Limitations acknowledged: The strongest BB-STATIC results require per-setting fraction selection; the noise edge only approximately separates signal from noise; evaluation is restricted to Pythia continued pretraining under the reported tuning protocol.

Related papers