Summary of "Hyperparameter Scaling Laws Across MoE Sparsity"
Summary (Overview)
-
Core Contribution: This paper establishes unified hyperparameter scaling laws for Mixture-of-Experts (MoE) models that explicitly incorporate sparsity (activation ratio ) as a predictive dimension, enabling reliable hyperparameter transfer across different sparsity levels.
-
Key Finding: At fixed sparsity, the optimal learning rate follows a power law in training compute , while the optimal batch size follows a power law in training tokens . The activation ratio enters both relationships as a multiplicative power-law factor.
-
Unified Formulation: The proposed scaling law takes the form , where , with and .
-
Scale of Experiments: 1,800 pre-training runs spanning six activated-parameter scales (10M to 324M), models up to 6B total non-embedding parameters, ~20 trillion tokens processed, at a cost of 200,000 equivalent H800 GPU-hours.
-
Validation: The laws jointly extrapolate to a held-out ultra-sparse MoE with 12B total parameters and , outperforming existing scaling laws and transferring across expert granularities.
Introduction and Theoretical Foundation
Background and Motivation
Mixture-of-Experts (MoE) models activate only a small subset of experts per token, expanding total parameter count without proportional increases in training compute. As model sizes grow, hyperparameters like learning rate (LR) and batch size (BS) become critical for training stability and final performance. However, exhaustive tuning at target scale is prohibitively expensive, motivating empirical scaling laws.
The Gap in Existing Work
Prior studies report conflicting findings on MoE hyperparameters:
- Some find hyperparameters transfer robustly between dense models and sparse MoEs (Wang et al., 2024; Li et al., 2025)
- Others observe MoEs favor larger batch sizes and lower learning rates (Ludziejewski et al., 2025; Tian et al., 2026)
Existing scaling laws fail to characterize how optima vary continuously with sparsity, particularly in the ultra-sparse regime down to .
Key Theoretical Insight
The paper demonstrates that neither activated parameter count nor total parameter count alone explains hyperparameter shifts across sparsity levels (Figure 2). The activation ratio:
must be modeled explicitly as an additional predictive dimension.
Methodology
Problem Formulation
The paper formalizes optimal hyperparameter selection via:
with the joint optimum defined as:
Experimental Setup
- Model Scales: Six activated-parameter scales (~10M to 324M), up to 6B total non-embedding parameters
- Sparsity Levels:
- Architecture: Hybrid linear-attention/MLA backbone with Muon optimizer
- Training: WSD schedule (1% warmup, stable peak, 10% exponential decay)
- Grid Search: Systematic variation of peak learning rate and global token batch size
- Optimality Definition: Near-optimal set defined as grid points within 0.1% of observed minimum loss
Candidate Variable Comparison
The study systematically compares candidate predictive variables:
- For : , , and (Figure 3a)
- For : , , and (Figure 4a)
Empirical Validation / Results
Base Scaling Laws at Fixed Sparsity
Optimal Learning Rate: Follows a power law in compute , insensitive to allocation at fixed :
Optimal Batch Size: Follows a power law in training tokens :
Fitted Coefficients
Table 2: Fitted coefficients of unified scaling laws
| Hyperparameter | Input variables | |||
|---|---|---|---|---|
| Learning rate | 0.8343 | -0.1385 | 0.1361 | |
| Batch size | 6.4765 | 0.5181 | -0.0841 |
The signs confirm: decreasing (more sparse) lowers optimal LR () and increases optimal BS ().
Functional Form Comparison
Table 3: Hyperparameter prediction errors for candidate families (LONO/LOAO errors)
| Candidate family | Functional form | p | BS error | LR error |
|---|---|---|---|---|
| Scale only | 2 | 0.295/0.271 | 0.316/0.355 | |
| Additive | 4 | 0.312/0.276 | 0.216/0.201 | |
| Log interaction | 4 | 0.307/0.238 | 0.205/0.150 | |
| Multiplicative (Ours) | 3 | 0.282/0.221 | 0.176/0.153 |
The multiplicative law was selected for its comparable performance with fewer parameters and simpler interpretation.
Joint Extrapolation to Held-Out Configuration
Table 4: Predictions on held-out target (, , , , )
| Method | Mode | Predicted LR | Predicted BS | Loss gap vs. ours |
|---|---|---|---|---|
| DeepSeek Law | Published | 7.22‰ | ||
| DeepSeek Law | Refitted | 1.37‰ | ||
| Step Law | Published | 9.02‰ | ||
| Step Law | Refitted | 1.03‰ | ||
| Ours | — | — |
Our prediction lies on the observed near-optimal loss plateau, demonstrating successful joint extrapolation.
Expert-Granularity Controls
Three controlled configurations at ~10M activated parameters:
- Reference: with
- Granularity control: — same , doubled expert counts, halved width
- Sparsity control: — same total experts,
Results show the activation-ratio relation transfers across expert granularities, and sparsity effects cannot be attributed solely to absolute expert counts.
Theoretical and Practical Implications
Theoretical Implications
-
Reconciliation of Conflicting Findings: The paper unifies prior contradictory results by showing that sparsity effects are multiplicative prefactor corrections to base power laws, explaining why some studies found transfer while others observed shifts.
-
Three Complementary Dimensions: Training compute , training tokens , and activation ratio provide three complementary predictors for MoE hyperparameters, extending the standard two-variable framework.
-
Gradient-Noise Interpretation: A simple gradient-noise model explains both trends: decreasing increases expert-side gradient noise (each expert receives only tokens per step), favoring larger global batches and lower learning rates.
Practical Implications
-
Transferable Prescription: The unified laws provide a practical recipe for selecting hyperparameters for ultra-sparse MoEs at scale without exhaustive tuning.
-
Cost Reduction: The laws enable hyperparameter selection for models with and 12B+ parameters based on smaller-scale experiments, potentially saving enormous tuning costs.
-
Sparsity as a Design Variable: The framework treats sparsity as an explicit predictor, allowing practitioners to anticipate hyperparameter adjustments when changing activation ratios.
Conclusion
Main Takeaways
This work systematically establishes that activation ratio is an essential predictive dimension for MoE hyperparameter scaling, alongside training compute and training tokens . The unified scaling form:
successfully:
- Reconciles conflicting findings in prior literature
- Outperforms alternative functional forms in grouped prediction
- Jointly extrapolates beyond fitting ranges in , , and
- Transfers across capacity-matched expert granularities
Future Directions
- Broader Validation: Testing across different architectures, optimizers, training schedules, routing mechanisms, and data mixtures
- Capability-Oriented Objectives: Extending from validation loss to downstream task performance
- Deeper Uncertainty Quantification: Addressing overlapping cross-validation intervals and limited extrapolation evidence
- Larger-Scale Confirmation: Repeating grid searches across multiple random seeds at lower cost
- Data and Efficiency Scaling: Developing complementary laws for data efficiency in ultra-sparse MoEs
The paper transforms sparsity from an architectural attribute into an explicit, quantifiable predictor for hyperparameter selection, providing a transferable prescription for training ultra-sparse MoEs at scale.
Related papers
- Do Tool Calls Execute as Intended? Measuring and Repairing Intent-Execution Correspondence in LLM Agents
Tool calls are silently altered by execution-path hops in 12% of production shell invocations, and the IntAct repair protocol recovers 79.2% of resulting failures.
- $σ$Transfer: Uncertainty Transfer from Small to Large Networks under $μ\mathrm{P}$
σTransfer enables zero-shot transfer of Laplace uncertainty estimates and prior precision from small to large neural networks, yielding up to 5000x speedups with negligible performance loss.
- Hybrid Latent Attention for Looped Language Models
Hybrid Latent Attention compresses looped language model KV caches 10.7x by having attention read compact latents directly, boosting decoding throughput up to 7.4x while retaining over 97% accuracy.