# Does Muon Need Fine-Grained Spectral Shaping?

> A single shared bulk-to-spike gain in MUON matches or beats fine-grained spectral shaping across 30 settings, showing useful spectral departures are surprisingly low-dimensional.

- **Source:** [arXiv](https://arxiv.org/abs/2610.07497)
- **Published:** 2026-10-10
- **Permalink:** https://picx.dev/p/PPn2X6
- **Whiteboard:** https://picx.dev/p/PPn2X6/image

## Summary

## Summary (Overview)

- **Core question**: Does MUON, a matrix-momentum optimizer for language models, need fine-grained spectral shaping of its update direction, or is a coarse two-band reweighting sufficient?
- **Key finding**: A single shared bulk-to-spike gain (BULKBOOST) captures at least as much benefit as fine-grained spectral profiles like FREON and SPECTRA, suggesting useful departures from MUON's flat spectrum are surprisingly low-dimensional.
- **Spectral diagnostics**: Approximately 93.7–97.1% of singular modes of MUON's momentum lie below a Marchenko–Pastur noise edge; these bulk directions are weak individually but collectively aligned with a reference gradient.
- **Empirical results**: Across 30 continued-pretraining settings (Pythia-14M to 410M, six corpora), two-band reweighting reduces final loss by 0.073–0.147% of pre-adaptation loss vs. MUON, outperforming FREON's 0.022% improvement.
- **Theoretical contribution**: Provides a first-order condition for when shifting weight toward the bulk lowers loss, and quantifies the fraction of maximal improvement rate that two bands can capture.

## Introduction and Theoretical Foundation

MUON combines current and past gradients into a matrix momentum $M_t$. Its SVD $M_t = U_t \Sigma_t V_t^\top$ decomposes the momentum into orthogonal directions weighted by singular values. The idealized polar update $Q_t = \mathrm{polar}(M_t) = U_t V_t^\top$ gives every singular direction unit gain—the **flat profile**.

Recent optimizers depart from this flat profile using spectral maps $U g(\Sigma) V^\top$ where gains vary across singular directions (FREON, SPECTRA, DEVA, and others). However, Shumaylov et al. (2026) showed that substantially different spectra, including random and inverted ones, can remain effective, raising the question: does MUON need fine-grained spectral shaping?

Two key observations motivate the two-band approach:

1. **Signal vs. strength are distinct**: A direction's singular value strength, its separation from sampling noise, and the benefit of increasing its update weight are different quantities. Small singular values can carry signal; large ones need not deserve more weight.

2. **Spectral bulk structure**: Under a Gaussian white-noise reference model, random matrix theory (Marchenko & Pastur, 1967; Baik et al., 2005) provides an asymptotic upper spectral edge. On Pythia-70M MUON trajectories, 93.7–97.1% of pooled singular modes lie in the bulk (below this edge), yet these weak directions collectively show positive signed alignment with an independently sampled reference gradient.

## Methodology

### MUON Setup and Notation

For a matrix parameter $W_t \in \mathbb{R}^{m \times n}$ with minibatch gradient $G_t$, the EMA buffer and Nesterov spectral input are:

$$B_t = \beta B_{t-1} + (1-\beta) G_t, \quad M_t = \beta B_t + (1-\beta) G_t \tag{1}$$

with $\beta \in [0,1)$ and $B_0 = 0$.

### Two-Band Partitions

BULKBOOST partitions singular directions into spike band $\mathcal{I}_{\mathrm{sp},t}$ and bulk band $\mathcal{I}_{\mathrm{bu},t}$:

- **BB-STATIC**: Fixed-rank rule with $r_t = \lceil p d \rceil$ for a fixed fraction $p$ (shared configuration uses $p = 5\%$)
- **BB-ADAPTIVE**: Noise-calibrated rule comparing singular values against threshold $\tau_t$:

$$\mathcal{I}_{\mathrm{sp},t} = \{i \in [d]: \sigma_{t,i} > \tau_t\}, \qquad \mathcal{I}_{\mathrm{bu},t} = \{i \in [d]: \sigma_{t,i} \leq \tau_t\} \tag{2}$$

### Norm-Matched Spectral Reweighting

The exact norm-matched direction for profile $w_t$ is:

$$D_t(w_t) = c_t(w_t) U_t \operatorname{diag}(w_t) V_t^\top, \qquad c_t(w_t) = \sqrt{\frac{d}{\sum_{i=1}^{d} w_{t,i}^2}} \tag{3}$$

BULKBOOST's two-level profile with bulk-to-spike gain $\rho$:

$$w_{t,i}^{(\rho)} = \begin{cases} 1, & i \in \mathcal{I}_{\mathrm{sp},t}, \\ \rho, & i \in \mathcal{I}_{\mathrm{bu},t}, \end{cases} \quad c_{\rho,r_t} := c_t(w_t^{(\rho)}) = \sqrt{\frac{d}{r_t + \rho^2(d - r_t)}} \tag{4}$$

### Noise-Edge Calibration (BB-ADAPTIVE)

Using two independent minibatch halves, the split-difference estimates noise:

$$\widehat{s}_t^2 = \frac{\|Z_t^{\mathrm{split}}\|_F^2}{mn}, \qquad \bar{s}_t^2 = \gamma \bar{s}_{t-1}^2 + (1-\gamma) \chi_t^2 \widehat{s}_t^2 \tag{6}$$

The threshold uses the propagated noise variance in the Nesterov input:

$$\nu_t^2(\beta) = (1-\beta)^2 \left[ (1+\beta)^2 + \frac{\beta^4(1-\beta^{2(t-1)})}{1-\beta^2} \right] \tag{7}$$

$$\tau_t = \hat{\tau}_t = \kappa_{\mathrm{edge}} \nu_t(\beta) \sqrt{\bar{s}_t^2}(\sqrt{m} + \sqrt{n}) \tag{8}$$

### SVD-Free Realization

An exact polar-plus-low-rank identity enables avoiding full SVD:

$$D_t^{\mathrm{svd}} = \begin{cases} c_{\rho,r_t}[\rho Q_t - (\rho-1) P_{t,\mathrm{sp}} Q_t], & m \leq n, \\ c_{\rho,r_t}[\rho Q_t - (\rho-1) Q_t R_{t,\mathrm{sp}}], & n < m. \end{cases} \tag{9}$$

where $P_{t,\mathrm{sp}} = U_{t,\mathrm{sp}} U_{t,\mathrm{sp}}^\top$ and $R_{t,\mathrm{sp}} = V_{t,\mathrm{sp}} V_{t,\mathrm{sp}}^\top$. Spike-subspace tracking uses the smaller Gram side with warm-started block power iterations.

## Empirical Validation / Results

### Theoretical Results

**Theorem 1 (Local spectral improvement)**: For a fixed partition with $r$ spike and $k = d-r$ bulk directions, with band means $\bar{u}_{\mathrm{sp}}$ and $\bar{u}_{\mathrm{bu}}$ and gap $\Delta = \bar{u}_{\mathrm{bu}} - \bar{u}_{\mathrm{sp}}$:

(i) Increasing $\rho$ from 1 gives strictly positive first-order improvement **iff** $\Delta > 0$.

(ii) Two bands reach fraction $R_{\mathrm{bt}} = \frac{\sqrt{rk/d}|\Delta|}{\|u_\perp\|_2} \in [0,1]$ of the best per-direction improvement rate.

**Theorem 2 (Exact polar-plus-low-rank decomposition)**: The unnormalized spectral update satisfies:

$$\widetilde{D}(a,\rho) = \rho Q + U_{\mathrm{sp}} \operatorname{diag}(a - \rho \mathbf{1}_r) V_{\mathrm{sp}}^\top, \quad \operatorname{rank}(\widetilde{D}(a,\rho) - \rho Q) \leq r \tag{14}$$

### Experimental Results

**Setup**: Continued pretraining of Pythia checkpoints (14M–410M) on six corpora (FineMath-4, FineMath-3, CodeParrot, Wiki-de, Wiki-zh, Wiki-ar), ~100M tokens per run, three seeds.

**Key results from Table 1** (final validation loss, mean ± std over three seeds):

| Model | Method | FineMath-4 | Wiki-zh | Wiki-ar |
|-------|--------|------------|---------|---------|
| 14M | MUON | 3.086 ± 0.002 | 2.745 ± 0.001 | 2.006 ± 0.000 |
| 14M | FREON | 3.083 ± 0.002 | 2.742 ± 0.001 | 2.008 ± 0.001 |
| 14M | BB-ADAPTIVE | 3.085 ± 0.002 | 2.737 ± 0.001 | 2.000 ± 0.001 |
| 70M | MUON | 2.551 ± 0.002 | 2.211 ± 0.005 | 1.702 ± 0.000 |
| 70M | FREON | 2.536 ± 0.001 | 2.200 ± 0.000 | 1.694 ± 0.001 |
| 70M | BB-ADAPTIVE | 2.545 ± 0.001 | 2.205 ± 0.004 | 1.690 ± 0.000 |
| 410M | MUON | 2.046 ± 0.005 | 1.879 ± 0.002 | 1.452 ± 0.001 |
| 410M | FREON | 2.043 ± 0.005 | 1.883 ± 0.001 | 1.460 ± 0.001 |
| 410M | BB-ADAPTIVE | 2.042 ± 0.005 | 1.875 ± 0.001 | 1.453 ± 0.002 |

**Summary statistics (Table 2)**:

| Method | Wins vs. MUON | Mean Δ vs. MUON (% of ℓ₀) | Tokens saved (%) |
|--------|---------------|--------------------------|------------------|
| FREON | 19/30 | -0.022 ± 0.010 | 7.78 |
| BB-STATIC(5%) | 23/30 | -0.073 ± 0.010 | 7.81 |
| BB-ADAPTIVE | 24/30 | -0.077 ± 0.010 | 7.99 |
| BB-STATIC† | 25/30 | -0.147 ± 0.021 | 10.00 |

**Key findings**:
- BB-STATIC(5%) beats FREON in 13 settings, loses in 13; BB-ADAPTIVE beats FREON in 17, loses in 10
- SPECTRA and DEVA trail MUON in all 30 settings
- BB-STATIC† beats BB-ADAPTIVE in 20 settings but requires loss-based fraction selection
- BB-ADAPTIVE matches BB-STATIC(5%) without fraction search (10 wins, 6 losses, 14 ties)
- Per-step runtimes: 1.36–1.50× MUON for BB-ADAPTIVE, 1.19–1.27× for static configurations

## Theoretical and Practical Implications

**Theoretical implications**:
- Near the flat profile, any norm-matched reweighting moves along some direction in the tangent space $\mathcal{T} = \{\delta \in \mathbb{R}^d: \mathbf{1}_d^\top \delta = 0\}$, with best rate $\|u_\perp\|_2$
- Two bands can only use the difference between band means; finer shaping adds at most the share $1 - R_{\mathrm{bt}}$ from within-band utility differences
- The polar-plus-low-rank identity (Theorem 2) shows the correction for two-band reweighting is inherently low-rank, enabling efficient SVD-free implementation

**Practical implications**:
- A single shared gain parameter $\rho$ suffices for competitive performance, eliminating the need for per-direction gain tuning
- BB-ADAPTIVE provides automatic noise-calibrated splitting without hyperparameter search, at moderate computational cost
- Token savings of 8–10% to reach MUON's final loss suggest meaningful efficiency gains in continued pretraining
- The approach transfers across model scales (14M–410M) and diverse corpora (math, code, multilingual text)

## Conclusion

The paper demonstrates that useful departures from MUON's flat spectral profile are surprisingly low-dimensional. A single bulk-to-spike gain captures at least as much benefit as fine-grained spectral profiles across 30 continued-pretraining settings. The key insights are:

1. **Spectral structure matters but coarsely**: The bulk (94–97% of modes below the noise edge) is weak individually but collectively useful, justifying a two-band treatment.

2. **Theory guides practice**: Theorem 1 provides a local test for when bulk emphasis helps; Theorem 2 enables efficient implementation without full SVD.

3. **Competitive performance**: Two-band reweighting matches or exceeds fine-grained baselines in aggregate, with BB-ADAPTIVE providing automatic split selection.

**Future directions**:
- Extending evaluation to other architectures and training regimes beyond Pythia continued pretraining
- Improving the noise-edge reference model beyond white Gaussian assumptions
- Exploring adaptive gain selection to replace the fixed $\rho = 4$ default
- Investigating whether the two-band insight transfers to other optimizer families

**Limitations acknowledged**: The strongest BB-STATIC results require per-setting fraction selection; the noise edge only approximately separates signal from noise; evaluation is restricted to Pythia continued pretraining under the reported tuning protocol.

---

_Markdown view of https://picx.dev/p/PPn2X6, served by PicX — AI-generated visual whiteboard summaries of research papers._
