# Hyperparameter Scaling Laws Across MoE Sparsity

> Unified hyperparameter scaling laws for MoE models treat activation ratio as a multiplicative power-law factor, enabling reliable learning rate and batch size transfer across sparsity levels.

- **Source:** [arXiv](https://arxiv.org/abs/2609.08690)
- **Published:** 2026-10-10
- **Permalink:** https://picx.dev/p/rXpBcP
- **Whiteboard:** https://picx.dev/p/rXpBcP/image

## Summary

# Summary of "Hyperparameter Scaling Laws Across MoE Sparsity"

## Summary (Overview)

- **Core Contribution**: This paper establishes unified hyperparameter scaling laws for Mixture-of-Experts (MoE) models that explicitly incorporate sparsity (activation ratio $A$) as a predictive dimension, enabling reliable hyperparameter transfer across different sparsity levels.

- **Key Finding**: At fixed sparsity, the optimal learning rate $\eta^*$ follows a power law in training compute $C = MD$, while the optimal batch size $B^*$ follows a power law in training tokens $D$. The activation ratio $A$ enters both relationships as a multiplicative power-law factor.

- **Unified Formulation**: The proposed scaling law takes the form $h^*(X, A) = k_h X^{\gamma_h} A^{\delta_h}$, where $(h, X) \in \{(\eta, C), (B, D)\}$, with $\delta_\eta > 0$ and $\delta_B < 0$.

- **Scale of Experiments**: 1,800 pre-training runs spanning six activated-parameter scales (10M to 324M), models up to 6B total non-embedding parameters, ~20 trillion tokens processed, at a cost of 200,000 equivalent H800 GPU-hours.

- **Validation**: The laws jointly extrapolate to a held-out ultra-sparse MoE with 12B total parameters and $A = 1/64$, outperforming existing scaling laws and transferring across expert granularities.

---

## Introduction and Theoretical Foundation

### Background and Motivation

Mixture-of-Experts (MoE) models activate only a small subset of experts per token, expanding total parameter count without proportional increases in training compute. As model sizes grow, hyperparameters like learning rate (LR) and batch size (BS) become critical for training stability and final performance. However, exhaustive tuning at target scale is prohibitively expensive, motivating empirical scaling laws.

### The Gap in Existing Work

Prior studies report **conflicting findings** on MoE hyperparameters:
- Some find hyperparameters transfer robustly between dense models and sparse MoEs (Wang et al., 2024; Li et al., 2025)
- Others observe MoEs favor larger batch sizes and lower learning rates (Ludziejewski et al., 2025; Tian et al., 2026)

Existing scaling laws fail to characterize how optima vary **continuously** with sparsity, particularly in the ultra-sparse regime down to $A = 1/64$.

### Key Theoretical Insight

The paper demonstrates that **neither activated parameter count $N$ nor total parameter count $N_{tot}$ alone** explains hyperparameter shifts across sparsity levels (Figure 2). The activation ratio:

$$A \equiv \frac{E_{act}}{E_{tot}}, \qquad C \equiv MD \tag{1}$$

must be modeled explicitly as an additional predictive dimension.

---

## Methodology

### Problem Formulation

The paper formalizes optimal hyperparameter selection via:

$$\mathcal{L}(\eta, B \mid N, N_{tot}, M, D, A) \tag{2}$$

with the joint optimum defined as:

$$(\eta^*, B^*) \equiv \arg\min_{\eta, B} \mathcal{L}(\eta, B \mid N, N_{tot}, M, D, A) \tag{3}$$

### Experimental Setup

- **Model Scales**: Six activated-parameter scales (~10M to 324M), up to 6B total non-embedding parameters
- **Sparsity Levels**: $A \in \{1, 1/4, 1/16, 1/32\}$
- **Architecture**: Hybrid linear-attention/MLA backbone with Muon optimizer
- **Training**: WSD schedule (1% warmup, stable peak, 10% exponential decay)
- **Grid Search**: Systematic variation of peak learning rate $\eta$ and global token batch size $B$
- **Optimality Definition**: Near-optimal set defined as grid points within 0.1% of observed minimum loss

### Candidate Variable Comparison

The study systematically compares candidate predictive variables:
- For $\eta^*$: $N$, $D$, and $C = MD$ (Figure 3a)
- For $B^*$: $N$, $C$, and $D$ (Figure 4a)

---

## Empirical Validation / Results

### Base Scaling Laws at Fixed Sparsity

**Optimal Learning Rate**: Follows a power law in compute $C = MD$, insensitive to $M/D$ allocation at fixed $C$:

$$\eta^*(C, A) = k_\eta C^{\gamma_\eta} A^{\delta_\eta} \tag{7}$$

**Optimal Batch Size**: Follows a power law in training tokens $D$:

$$B^*(D, A) = k_B D^{\gamma_B} A^{\delta_B} \tag{8}$$

### Fitted Coefficients

**Table 2: Fitted coefficients of unified scaling laws**

| Hyperparameter $h$ | Input variables | $k_h$ | $\gamma_h$ | $\delta_h$ |
|---|---|---|---|---|
| Learning rate $\eta$ | $(C, A)$ | 0.8343 | -0.1385 | 0.1361 |
| Batch size $B$ | $(D, A)$ | 6.4765 | 0.5181 | -0.0841 |

The signs confirm: decreasing $A$ (more sparse) lowers optimal LR ($\delta_\eta > 0$) and increases optimal BS ($\delta_B < 0$).

### Functional Form Comparison

**Table 3: Hyperparameter prediction errors for candidate families** (LONO/LOAO errors)

| Candidate family | Functional form | p | BS error | LR error |
|---|---|---|---|---|
| Scale only | $kX^{\gamma}$ | 2 | 0.295/0.271 | 0.316/0.355 |
| Additive | $k_X X^{\gamma} + k_A A^{\delta}$ | 4 | 0.312/0.276 | 0.216/0.201 |
| Log interaction | $kX^{\beta_X + \beta_{XA}\log_2 A} A^{\beta_A}$ | 4 | 0.307/0.238 | 0.205/0.150 |
| **Multiplicative (Ours)** | $kX^{\gamma}A^{\delta}$ | 3 | **0.282/0.221** | **0.176/0.153** |

The multiplicative law was selected for its comparable performance with fewer parameters and simpler interpretation.

### Joint Extrapolation to Held-Out Configuration

**Table 4: Predictions on held-out target** ($N = 324M$, $N_{tot} = 12B$, $D = 159B$, $A = 1/64$, $C = 3 \times 10^{20}$)

| Method | Mode | Predicted LR | Predicted BS | Loss gap vs. ours |
|---|---|---|---|---|
| DeepSeek Law | Published | $8.59 \times 10^{-4}$ | $1.46 \times 10^6$ | 7.22‰ |
| DeepSeek Law | Refitted | $8.11 \times 10^{-4}$ | $2.73 \times 10^6$ | 1.37‰ |
| Step Law | Published | $3.20 \times 10^{-4}$ | $1.44 \times 10^6$ | 9.02‰ |
| Step Law | Refitted | $8.99 \times 10^{-4}$ | $5.13 \times 10^6$ | 1.03‰ |
| **Ours** | — | $6.92 \times 10^{-4}$ | $5.84 \times 10^6$ | — |

Our prediction lies on the observed near-optimal loss plateau, demonstrating successful joint extrapolation.

### Expert-Granularity Controls

Three controlled configurations at ~10M activated parameters:
- **Reference**: $(E_{act}, E_{tot}, h_{MoE}) = (2, 64, 384)$ with $A = 1/32$
- **Granularity control**: $(4, 128, 192)$ — same $A$, doubled expert counts, halved width
- **Sparsity control**: $(4, 64, 384)$ — same total experts, $A = 1/16$

Results show the activation-ratio relation transfers across expert granularities, and sparsity effects cannot be attributed solely to absolute expert counts.

---

## Theoretical and Practical Implications

### Theoretical Implications

1. **Reconciliation of Conflicting Findings**: The paper unifies prior contradictory results by showing that sparsity effects are multiplicative prefactor corrections to base power laws, explaining why some studies found transfer while others observed shifts.

2. **Three Complementary Dimensions**: Training compute $C$, training tokens $D$, and activation ratio $A$ provide three complementary predictors for MoE hyperparameters, extending the standard two-variable framework.

3. **Gradient-Noise Interpretation**: A simple gradient-noise model explains both trends: decreasing $A$ increases expert-side gradient noise (each expert receives only $AB$ tokens per step), favoring larger global batches and lower learning rates.

### Practical Implications

1. **Transferable Prescription**: The unified laws provide a practical recipe for selecting hyperparameters for ultra-sparse MoEs at scale without exhaustive tuning.

2. **Cost Reduction**: The laws enable hyperparameter selection for models with $A = 1/64$ and 12B+ parameters based on smaller-scale experiments, potentially saving enormous tuning costs.

3. **Sparsity as a Design Variable**: The framework treats sparsity as an explicit predictor, allowing practitioners to anticipate hyperparameter adjustments when changing activation ratios.

---

## Conclusion

### Main Takeaways

This work systematically establishes that **activation ratio $A$ is an essential predictive dimension** for MoE hyperparameter scaling, alongside training compute $C$ and training tokens $D$. The unified scaling form:

$$h^*(X, A) = k_h X^{\gamma_h} A^{\delta_h}$$

successfully:
- Reconciles conflicting findings in prior literature
- Outperforms alternative functional forms in grouped prediction
- Jointly extrapolates beyond fitting ranges in $A$, $D$, and $C$
- Transfers across capacity-matched expert granularities

### Future Directions

1. **Broader Validation**: Testing across different architectures, optimizers, training schedules, routing mechanisms, and data mixtures
2. **Capability-Oriented Objectives**: Extending from validation loss to downstream task performance
3. **Deeper Uncertainty Quantification**: Addressing overlapping cross-validation intervals and limited extrapolation evidence
4. **Larger-Scale Confirmation**: Repeating grid searches across multiple random seeds at lower cost
5. **Data and Efficiency Scaling**: Developing complementary laws for data efficiency in ultra-sparse MoEs

The paper transforms sparsity from an architectural attribute into an explicit, quantifiable predictor for hyperparameter selection, providing a transferable prescription for training ultra-sparse MoEs at scale.

---

_Markdown view of https://picx.dev/p/rXpBcP, served by PicX — AI-generated visual whiteboard summaries of research papers._
