Summary (Overview)
- Challenge to established practice: The paper demonstrates that batch-size scaling rules for optimizers do not transfer consistently across training settings, and that optimizer rankings can reverse across batch sizes even after extensive hyperparameter tuning.
- Key empirical finding: In language model pretraining with the Modded-NanoGPT benchmark, SOAP performs best at small batch sizes (128K–1M tokens) while Shampoo wins at the largest batch size (2M tokens), with crossovers also observed between Muon and Shampoo.
- Theoretical contribution: The paper proves that gradient noise drives a bias–variance tradeoff in preconditioning that can reverse optimizer rankings across batch sizes (Theorem 5.1), even under optimal learning rate and momentum tuning.
- Direction-dependent scaling: The optimal batch-size scaling exponent shifts from approximately square-root toward linear as a direction's curvature-to-noise ratio (CNR) increases, explaining why single-exponent scaling rules fail.
- Practical validation: Retaining small-batch updates in only the sharpest 0.001% of directions removes up to 59.5% of the loss gap when increasing batch size 16×, while a random subspace of the same size barely helps.
Introduction and Theoretical Foundation
Background and Motivation
Larger batches can increase training throughput but often require hyperparameter adjustments to retain optimization performance. Batch-size scaling rules promise to prescribe these adjustments without additional tuning, derived analytically and intended to apply broadly across training settings. However, different derivations yield conflicting prescriptions for the same optimizer.
The paper poses two key questions:
- Q1: Can a hyperparameter scaling rule consistently work in different training settings?
- Q2: Do optimizer rankings persist across batch sizes after extensive retuning?
Theoretical Frameworks for Scaling Rules
Invariant-preservation view: Adjusts hyperparameters to preserve a training property such as the parameter trajectory or data remembered by moving averages. For SGD, the SDE approximation gives:
Preserving gives the linear scaling rule .
Bound-minimization view: Starts from convergence bounds on suboptimality or stationarity, minimizing over scaled hyperparameters with fixed sample-time budget.
Optimizer Definitions
MuonW–AdamW (Definition 1): Combines Muon with decoupled weight decay for matrix parameters and AdamW for embeddings. The update rules are:
Noisy Quadratic Model (NQM) (Definition 4): For with curvature and noise variance :
Curvature-to-noise ratio (CNR) (Definition 5):
Methodology
Deriving Scaling Rules for Muon
Candidate 1 (SDE matching) (Definition 2): Matching the coupled SDE for parameters and momentum gives:
Candidate 2 (Bound minimization) (Definition 3): Minimizing the Frank–Wolfe gap bound gives:
Empirical Sweep Protocol
- Tested complete scaling rules in language modeling and on CIFAR-5M
- Reference batches: 128K tokens (language modeling) and 256 examples (CIFAR-5M)
- Selection criterion: mean batch-size regret across target batches
Optimizer Benchmarking
- Setting: Modded-NanoGPT optimization benchmark with batch sizes B ∈ {128K, 512K, 1M, 2M} tokens, ~1.7B token budget
- Optimizers compared: Adam, Lion, Muon, SOAP, Shampoo
- Independent tuning at every batch size via coordinate descent
- Two weight-norm control mechanisms: decoupled weight decay (W) and HyperBall (H)
Directional Branching Experiment
- Branched from checkpoints of a tuned MuonW run at B = 128K tokens
- Emulated 2M-token training by averaging 16 gradients (κ = 16)
- Held subspaces: top-16 and top-768 sharpest directions of a preconditioned Hessian proxy, plus a random 768-dim baseline
Empirical Validation / Results
Scaling Rule Failure (Figure 2)
- Language modeling: The empirical LLM rule (fixed learning rates, linear weight-decay scaling) stays closest to full retuning; both theoretical rules fall behind no scaling from 1M tokens onward
- CIFAR-5M: The bound-minimization rule stays within 0.006 nats of full retuning, while the empirical LLM rule does no better than no scaling
- Best rule depends on setting: Language modeling favors fixed learning rates with linear matrix weight-decay scaling; CIFAR-5M favors square-root learning-rate scaling with fixed weight decay
Optimizer Crossover (Figure 1)
- SOAP is best from 128K to 1M tokens (ahead by up to 0.012 nats), but falls to third at 2M, 0.021 nats behind Shampoo
- Muon beats Shampoo by 0.0035 nats at 512K, but Shampoo beats Muon by 0.0037 nats at 2M
- Same crossover pattern observed under HyperBall norm control (Figure 3)
NQM Analysis
Scaling depends on CNR (Figure 4d): Increasing CNR shifts tuned learning rates from approximately square-root toward linear scaling.
Theoretical result (Theorem 5.1): For every sufficiently large sample budget T, there exists a two-dimensional NQM where optimally tuned SGD outperforms optimally tuned Newton at B = 1, but the ranking reverses at B = T. This holds both without momentum and with optimally tuned momentum.
Mechanism: The expected sign of the momentum average at fixed parameter value w is:
When noise dominates, this grows as (motivating square-root scaling); when signal dominates, it saturates at 1 (motivating linear scaling).
Directional Branching Results (Figure 6)
| Held Subspace | Fraction of Directions | Loss Gap Removed |
|---|---|---|
| Top-16 sharpest | 0.00002% | Up to 27% |
| Top-768 sharpest | 0.001% | 35–59.5% |
| Random 768-dim | 0.001% | ≤ 3% |
The effect disappears for branches started after learning-rate decay begins (steps 11,000 and 12,000).
Theoretical and Practical Implications
Key Theoretical Insights
-
Direction-dependent scaling: The optimal scaling exponent varies across directions based on CNR, making single-exponent scaling rules fundamentally limited. The shift from square-root to linear scaling reflects the transition from noise-dominated to signal-dominated dynamics.
-
Bias–variance tradeoff in preconditioning: Newton's method (and by extension adaptive preconditioners) trades faster reduction of initialization bias for greater variance from gradient noise amplification in flat directions. This tradeoff is batch-size-dependent, explaining why optimizer rankings can reverse.
-
Momentum's batch-dependent role: Keeping momentum coefficients fixed is generally preferable in sweeps. Momentum becomes useful at large batches by increasing the maximum tolerable learning rate in sharp directions (consistent with river-valley theory).
Practical Implications
- Benchmarking: Single-batch-size comparisons are insufficient to establish optimizer efficacy; extensive tuning across several batch sizes is recommended
- Scaling rules: The promise of setting-agnostic scaling rules is not fulfilled—rules must account for training setting, target batch size, and direction-dependent CNR
- Optimizer design: Must account for how batch size changes the benefits and costs of adaptation, not just efficiently estimate landscape statistics
Conclusion
The paper challenges established practices in optimizer evaluation and design by demonstrating that:
- No principled scaling rule for Muon works consistently across training settings—the best rule depends on the task, target batch size, and reference batch size
- Optimizer rankings reverse across batch sizes even after extensive hyperparameter tuning, with SOAP winning at small batches and Shampoo at large batches
- The root cause is direction-dependent: the optimal scaling exponent shifts from square-root to linear as a direction's curvature-to-noise ratio increases, driven by a bias–variance tradeoff in preconditioning
The empirical validation with directional branching confirms that the local batch-size penalty in language model pretraining is concentrated in the sharpest directions—retaining small-batch updates in just 0.001% of directions removes up to 59.5% of the loss gap.
Future directions: The paper calls for optimizer benchmarking across multiple batch sizes, scaling rules that account for CNR variation across directions, and further exploration of how gradient noise anisotropy affects direction selection in batch-size adaptation.
Related papers
- $σ$Transfer: Uncertainty Transfer from Small to Large Networks under $μ\mathrm{P}$
σTransfer enables zero-shot transfer of Laplace uncertainty estimates and prior precision from small to large neural networks, yielding up to 5000x speedups with negligible performance loss.
- Memory-Efficient Expert Routing for Distributed MoE Training
RelayMoE replaces all-to-all MoE dispatch with ring-based local computation, cutting peak memory up to 7.6x and speeding training throughput by 2x.
- Normalize-Then-Precondition: A Hierarchical Approach to Marginal Scale and Interaction Geometry for LLM Training
NormPre separates LLM update normalization from spectral preconditioning, consistently beating AdamW, Muon, and MANO in pretraining while cutting optimizer latency by up to 67%.