Summary of "Hyperparameter Scaling Laws Across MoE Sparsity"

Summary (Overview)

  • Core Contribution: This paper establishes unified hyperparameter scaling laws for Mixture-of-Experts (MoE) models that explicitly incorporate sparsity (activation ratio AA) as a predictive dimension, enabling reliable hyperparameter transfer across different sparsity levels.

  • Key Finding: At fixed sparsity, the optimal learning rate η∗\eta^* follows a power law in training compute C=MDC = MD, while the optimal batch size B∗B^* follows a power law in training tokens DD. The activation ratio AA enters both relationships as a multiplicative power-law factor.

  • Unified Formulation: The proposed scaling law takes the form h∗(X,A)=khXγhAδhh^*(X, A) = k_h X^{\gamma_h} A^{\delta_h}, where (h,X)∈{(η,C),(B,D)}(h, X) \in \{(\eta, C), (B, D)\}, with δη>0\delta_\eta > 0 and δB<0\delta_B < 0.

  • Scale of Experiments: 1,800 pre-training runs spanning six activated-parameter scales (10M to 324M), models up to 6B total non-embedding parameters, ~20 trillion tokens processed, at a cost of 200,000 equivalent H800 GPU-hours.

  • Validation: The laws jointly extrapolate to a held-out ultra-sparse MoE with 12B total parameters and A=1/64A = 1/64, outperforming existing scaling laws and transferring across expert granularities.


Introduction and Theoretical Foundation

Background and Motivation

Mixture-of-Experts (MoE) models activate only a small subset of experts per token, expanding total parameter count without proportional increases in training compute. As model sizes grow, hyperparameters like learning rate (LR) and batch size (BS) become critical for training stability and final performance. However, exhaustive tuning at target scale is prohibitively expensive, motivating empirical scaling laws.

The Gap in Existing Work

Prior studies report conflicting findings on MoE hyperparameters:

  • Some find hyperparameters transfer robustly between dense models and sparse MoEs (Wang et al., 2024; Li et al., 2025)
  • Others observe MoEs favor larger batch sizes and lower learning rates (Ludziejewski et al., 2025; Tian et al., 2026)

Existing scaling laws fail to characterize how optima vary continuously with sparsity, particularly in the ultra-sparse regime down to A=1/64A = 1/64.

Key Theoretical Insight

The paper demonstrates that neither activated parameter count NN nor total parameter count NtotN_{tot} alone explains hyperparameter shifts across sparsity levels (Figure 2). The activation ratio:

A≡EactEtot,C≡MD(1)A \equiv \frac{E_{act}}{E_{tot}}, \qquad C \equiv MD \tag{1}

must be modeled explicitly as an additional predictive dimension.


Methodology

Problem Formulation

The paper formalizes optimal hyperparameter selection via:

L(η,B∣N,Ntot,M,D,A)(2)\mathcal{L}(\eta, B \mid N, N_{tot}, M, D, A) \tag{2}

with the joint optimum defined as:

(η∗,B∗)≡arg⁡min⁡η,BL(η,B∣N,Ntot,M,D,A)(3)(\eta^*, B^*) \equiv \arg\min_{\eta, B} \mathcal{L}(\eta, B \mid N, N_{tot}, M, D, A) \tag{3}

Experimental Setup

  • Model Scales: Six activated-parameter scales (~10M to 324M), up to 6B total non-embedding parameters
  • Sparsity Levels: A∈{1,1/4,1/16,1/32}A \in \{1, 1/4, 1/16, 1/32\}
  • Architecture: Hybrid linear-attention/MLA backbone with Muon optimizer
  • Training: WSD schedule (1% warmup, stable peak, 10% exponential decay)
  • Grid Search: Systematic variation of peak learning rate η\eta and global token batch size BB
  • Optimality Definition: Near-optimal set defined as grid points within 0.1% of observed minimum loss

Candidate Variable Comparison

The study systematically compares candidate predictive variables:

  • For η∗\eta^*: NN, DD, and C=MDC = MD (Figure 3a)
  • For B∗B^*: NN, CC, and DD (Figure 4a)

Empirical Validation / Results

Base Scaling Laws at Fixed Sparsity

Optimal Learning Rate: Follows a power law in compute C=MDC = MD, insensitive to M/DM/D allocation at fixed CC:

η∗(C,A)=kηCγηAδη(7)\eta^*(C, A) = k_\eta C^{\gamma_\eta} A^{\delta_\eta} \tag{7}

Optimal Batch Size: Follows a power law in training tokens DD:

B∗(D,A)=kBDγBAδB(8)B^*(D, A) = k_B D^{\gamma_B} A^{\delta_B} \tag{8}

Fitted Coefficients

Table 2: Fitted coefficients of unified scaling laws

Hyperparameter hhInput variableskhk_hγh\gamma_hδh\delta_h
Learning rate η\eta(C,A)(C, A)0.8343-0.13850.1361
Batch size BB(D,A)(D, A)6.47650.5181-0.0841

The signs confirm: decreasing AA (more sparse) lowers optimal LR (δη>0\delta_\eta > 0) and increases optimal BS (δB<0\delta_B < 0).

Functional Form Comparison

Table 3: Hyperparameter prediction errors for candidate families (LONO/LOAO errors)

Candidate familyFunctional formpBS errorLR error
Scale onlykXγkX^{\gamma}20.295/0.2710.316/0.355
AdditivekXXγ+kAAδk_X X^{\gamma} + k_A A^{\delta}40.312/0.2760.216/0.201
Log interactionkXβX+βXAlog⁡2AAβAkX^{\beta_X + \beta_{XA}\log_2 A} A^{\beta_A}40.307/0.2380.205/0.150
Multiplicative (Ours)kXγAδkX^{\gamma}A^{\delta}30.282/0.2210.176/0.153

The multiplicative law was selected for its comparable performance with fewer parameters and simpler interpretation.

Joint Extrapolation to Held-Out Configuration

Table 4: Predictions on held-out target (N=324MN = 324M, Ntot=12BN_{tot} = 12B, D=159BD = 159B, A=1/64A = 1/64, C=3×1020C = 3 \times 10^{20})

MethodModePredicted LRPredicted BSLoss gap vs. ours
DeepSeek LawPublished8.59×10−48.59 \times 10^{-4}1.46×1061.46 \times 10^67.22‰
DeepSeek LawRefitted8.11×10−48.11 \times 10^{-4}2.73×1062.73 \times 10^61.37‰
Step LawPublished3.20×10−43.20 \times 10^{-4}1.44×1061.44 \times 10^69.02‰
Step LawRefitted8.99×10−48.99 \times 10^{-4}5.13×1065.13 \times 10^61.03‰
Ours—6.92×10−46.92 \times 10^{-4}5.84×1065.84 \times 10^6—

Our prediction lies on the observed near-optimal loss plateau, demonstrating successful joint extrapolation.

Expert-Granularity Controls

Three controlled configurations at ~10M activated parameters:

  • Reference: (Eact,Etot,hMoE)=(2,64,384)(E_{act}, E_{tot}, h_{MoE}) = (2, 64, 384) with A=1/32A = 1/32
  • Granularity control: (4,128,192)(4, 128, 192) — same AA, doubled expert counts, halved width
  • Sparsity control: (4,64,384)(4, 64, 384) — same total experts, A=1/16A = 1/16

Results show the activation-ratio relation transfers across expert granularities, and sparsity effects cannot be attributed solely to absolute expert counts.


Theoretical and Practical Implications

Theoretical Implications

  1. Reconciliation of Conflicting Findings: The paper unifies prior contradictory results by showing that sparsity effects are multiplicative prefactor corrections to base power laws, explaining why some studies found transfer while others observed shifts.

  2. Three Complementary Dimensions: Training compute CC, training tokens DD, and activation ratio AA provide three complementary predictors for MoE hyperparameters, extending the standard two-variable framework.

  3. Gradient-Noise Interpretation: A simple gradient-noise model explains both trends: decreasing AA increases expert-side gradient noise (each expert receives only ABAB tokens per step), favoring larger global batches and lower learning rates.

Practical Implications

  1. Transferable Prescription: The unified laws provide a practical recipe for selecting hyperparameters for ultra-sparse MoEs at scale without exhaustive tuning.

  2. Cost Reduction: The laws enable hyperparameter selection for models with A=1/64A = 1/64 and 12B+ parameters based on smaller-scale experiments, potentially saving enormous tuning costs.

  3. Sparsity as a Design Variable: The framework treats sparsity as an explicit predictor, allowing practitioners to anticipate hyperparameter adjustments when changing activation ratios.


Conclusion

Main Takeaways

This work systematically establishes that activation ratio AA is an essential predictive dimension for MoE hyperparameter scaling, alongside training compute CC and training tokens DD. The unified scaling form:

h∗(X,A)=khXγhAδhh^*(X, A) = k_h X^{\gamma_h} A^{\delta_h}

successfully:

  • Reconciles conflicting findings in prior literature
  • Outperforms alternative functional forms in grouped prediction
  • Jointly extrapolates beyond fitting ranges in AA, DD, and CC
  • Transfers across capacity-matched expert granularities

Future Directions

  1. Broader Validation: Testing across different architectures, optimizers, training schedules, routing mechanisms, and data mixtures
  2. Capability-Oriented Objectives: Extending from validation loss to downstream task performance
  3. Deeper Uncertainty Quantification: Addressing overlapping cross-validation intervals and limited extrapolation evidence
  4. Larger-Scale Confirmation: Repeating grid searches across multiple random seeds at lower cost
  5. Data and Efficiency Scaling: Developing complementary laws for data efficiency in ultra-sparse MoEs

The paper transforms sparsity from an architectural attribute into an explicit, quantifiable predictor for hyperparameter selection, providing a transferable prescription for training ultra-sparse MoEs at scale.

Related papers