# ID Balancing: Stable Training of Extremely Sparse MoE via PID-Based Load Control

> ID Balancing, a magnitude-aware integral-derivative controller with worsening-gated derivative and bias re-centering, reduces worst-case expert overload by over 50% versus best baselines in sparse MoE training.

- **Source:** [arXiv](https://arxiv.org/abs/2609.39137)
- **Published:** 2026-10-03
- **Permalink:** https://picx.dev/p/DiQXOy
- **Whiteboard:** https://picx.dev/p/DiQXOy/image

## Summary

## Summary (Overview)

- **Unified PID framework**: The paper proposes a generalized Proportional–Integral–Derivative (PID) control view of auxiliary-loss-free MoE load balancing, showing that DeepSeek's loss-free method acts as a fixed-step integral controller while Kimi K3's Quantile Balancing functions as a generalized proportional controller.
- **ID Balancing method**: Introduces an Integral–Derivative controller combining a magnitude-aware integral term with a worsening-gated derivative term, requiring only $O(E)$ token-count feedback per layer with no balancing gradient added to the training objective.
- **Superior load control**: Reduces worst-case backbone MaxVio by over 50% and training-average backbone MinVio by over 12% relative to the best baselines in the Top-3-of-768 setting, with advantages growing as sparsity increases.
- **Scale robustness**: Under fixed Top-10-of-768 routing, worst-case backbone MaxVio remains nearly unchanged as total parameters increase from 18.9B to 69.9B (approximately 89.6% lower than the auxiliary-loss baseline).
- **Competitive performance**: Maintains competitive language-modeling loss and downstream performance across nine benchmarks while achieving effective load control.

## Introduction and Theoretical Foundation

### Background and Motivation

Mixture-of-Experts (MoE) enables efficient scaling of large language models through sparse expert activation. For each token, a router selects only $K$ experts from a pool of $E$. With $K$ fixed, adding experts increases total capacity without increasing per-token computation. However, as the active fraction $K/E$ decreases, **expert load imbalance becomes more severe**:
- Overloaded experts slow down MoE computation
- Underloaded experts receive insufficient training, leaving model capacity unused

### Control-Theoretic Formulation

The paper formulates MoE load balancing as a control problem. For an MoE layer with $E$ routed experts and a Top-$K$ router:

- Router logits: $z = W_r x$ where $W_r \in \mathbb{R}^{E \times d}$ is the router weight matrix
- Expert scores: $s_i = \sigma(z_i)$ via activation function $\sigma$
- Expert selection with bias: $T(x) = \text{TopK}_{i \in [E]}(s_i + b_i)$
- Normalized load error: $e_i^{(t)} = \frac{\bar{n}^{(t)} - n_i^{(t)}}{\bar{n}^{(t)}}$ where $\bar{n}^{(t)} = \frac{1}{E}\sum_{i=1}^{E} n_i^{(t)}$

The standard PID controller is given by:

$$u_t = K_p e_t + K_i \sum_{\tau \leq t} e_\tau + K_d(e_t - e_{t-1})$$

where $K_p, K_i, K_d \geq 0$ are the gains of the proportional, integral, and derivative branches.

### Existing Methods Through a PID Lens

Under the generalized PID view (where temporal roles define branches, allowing nonlinear transformations):

**DeepSeek's loss-free method** (fixed-step integral control):
$$b^{(t+1)} = b^{(t)} + \eta \cdot \text{sign}(e^{(t)}) = \eta \sum_{\tau \leq t} \text{sign}(e^{(\tau)}) = I^{(t)}$$

The sign operation discards error magnitude—every nonzero error receives the same correction $\eta$.

**Quantile Balancing** (generalized proportional control):
$$b^{(t+1)} = Q(S^{(t)}, b^{(t)}) = P^{(t)}$$

where $Q_j(S^{(t)}, b^{(t)}) = -\text{quantile}_{1-K/E}\left(\{S_{i,j}^{(t)} - \alpha_i^{(t)}\}_{i=1}^{m}\right)$

This directly responds to batch-dependent targets without explicitly accumulating past errors.

## Methodology

### ID Balancing Algorithm

The proposed ID Balancing combines two key components:

**1. Magnitude-aware integral term:**
$$b^{(t+1)}_i = b^{(t)}_i + K_i e^{(t)}_i$$

The correction grows with $|e^{(t)}_i|$ and shrinks near balance. Unlike the sign-based update, this preserves zero-mean bias when initialized at $b^{(0)} = 0$ since $\sum_i e^{(t)}_i = 0$.

**2. Worsening-gated derivative term:**
$$\Delta e^{(t)}_i = e^{(t)}_i - e^{(t-1)}_i$$
$$g^{(t)}_i = \mathbb{1}[e^{(t-1)}_i \cdot \Delta e^{(t)}_i > 0]$$

The gate opens only when the imbalance worsens (error and its change have the same sign), providing extra correction only during deterioration.

**Complete update with re-centering:**
$$\tilde{b}^{(t+1)}_i = b^{(t)}_i + K_i e^{(t)}_i + K_d g^{(t)}_i \Delta e^{(t)}_i$$
$$b^{(t+1)}_i = \tilde{b}^{(t+1)}_i - \frac{1}{E}\sum_{j=1}^{E} \tilde{b}^{(t+1)}_j$$

The zero-mean centering removes common bias drift without changing routing decisions (adding the same constant to all biases leaves the ranking unchanged).

### Algorithm Pseudocode

```
Algorithm 1: ID Balancing bias update, one MoE layer, per training iteration
Require: expert count E, active experts K, integral gain K_i, derivative gain K_d
1: b_i ← 0, e_prev_i ← 0 for all i ∈ [E]
2: for each training iteration do
3:   route the batch by T(x) = TopK_{i∈[E]}(s_i + b_i)
4:   n_i ← number of tokens assigned to expert i
5:   n̄ ← (1/E) Σ_{i=1}^{E} n_i
6:   for each expert i ∈ [E] do
7:     e_i ← (n̄ - n_i) / n̄
8:     Δe_i ← e_i - e_prev_i
9:     g_i ← 1[e_prev_i · Δe_i > 0]
10:    b_i ← b_i + K_i e_i + K_d g_i Δe_i
11:    e_prev_i ← e_i
12:  end for
13:  b_i ← b_i - (1/E) Σ_{j=1}^{E} b_j for all i ∈ [E]
14: end for
```

### Experimental Setup

- **Architecture**: Qwen-3.8-Next with 20 Transformer layers, one multi-token prediction (MTP) module, 768 routed experts per MoE layer
- **Model sizes**: 18.9B total parameters (0.87–1.03B active), 24.8B, and 69.9B total parameters
- **Routing settings**: Top-10, Top-5, and Top-3 over 768 experts
- **Training**: 120B tokens (main comparison), 560B tokens (downstream evaluation), global batch size 1024
- **Learning rate**: Peak 2.54 × 10⁻³ with cosine decay to 3 × 10⁻⁵; stress test at constant 5.86 × 10⁻³ (2.3× standard peak)
- **Default gains**: $K_i = K_d = 6 \times 10^{-3}$ (reduced to $6 \times 10^{-6}$ for continued pretraining)

### Evaluation Metrics

$$\text{MaxVio}^{(t)} = \max_i \frac{n_i^{(t)} - \bar{n}^{(t)}}{\bar{n}^{(t)}}, \quad \text{MinVio}^{(t)} = \frac{\bar{n}^{(t)} - \min_i n_i^{(t)}}{\bar{n}^{(t)}}$$

- **MaxVio**: largest relative overload (bounded by $E/K - 1$)
- **MinVio**: largest relative deficit (in [0, 1], reaching 1 when an expert receives no tokens)

## Empirical Validation / Results

### Main Results (Table 1): Load Control Across Routing Sparsities

| Method | Act./Total | LM loss | Backbone MaxVio (Worst) | Backbone MinVio (Avg.) |
|--------|-----------|---------|------------------------|------------------------|
| **Top-3 of 768 experts** | 0.87B/18.9B | | | |
| Auxiliary loss | | 1.7527 | 70.76 | 0.6232 |
| DeepSeek loss-free | | 1.7426 | 211.19 | 0.7441 |
| Quantile Balancing | | 1.7432 | 31.11 | 0.6373 |
| **ID Balancing** | | **1.7426** | **15.33** | **0.5480** |
| **Top-5 of 768 experts** | 0.91B/18.9B | | | |
| Auxiliary loss | | 1.7378 | 50.48 | 0.5836 |
| DeepSeek loss-free | | 1.7280 | 115.82 | 0.6725 |
| Quantile Balancing | | 1.7303 | 21.16 | 0.5232 |
| **ID Balancing** | | **1.7299** | **10.33** | **0.4601** |
| **Top-10 of 768 experts** | 1.03B/18.9B | | | |
| Auxiliary loss | | 1.7180 | 16.42 | 0.5446 |
| DeepSeek loss-free | | 1.7121 | 69.07 | 0.5571 |
| Quantile Balancing | | 1.7133 | 10.13 | 0.3914 |
| **ID Balancing** | | **1.7118** | **8.36** | **0.3800** |

**Key findings**: ID Balancing achieves the lowest Worst and Last1k MaxVio, as well as the lowest Last1k and Avg. MinVio across all three routing settings, with LM loss gaps under 0.0023 among bias-based methods.

### Scaling Results (Figure 1)

- **Sparsity scaling**: As routing becomes sparser (Top-10 → Top-3 of 768), ID Balancing's worst-case MaxVio grows from 8.36 to 15.33, while DeepSeek's grows from 69.07 to 211.19
- **Model size scaling**: Under fixed Top-10-of-768 routing, ID Balancing's worst-case MaxVio remains nearly unchanged (8.36 → 7.79) as total parameters increase from 18.9B to 69.9B, approximately 89.6% lower than the auxiliary-loss baseline

### Higher Learning Rate Stress Test (Table 3)

At constant learning rate 5.86 × 10⁻³ (2.3× standard peak):

| Method | LM loss | Backbone MaxVio (Worst) | Backbone MinVio (Avg.) |
|--------|---------|------------------------|------------------------|
| Auxiliary loss | 2.0099 | 20.89 | 0.8458 |
| DeepSeek loss-free | 2.0031 | 72.54 | 0.8269 |
| Quantile Balancing | 2.0009 | 2.96 | 0.3938 |
| **ID Balancing** | **2.0007** | 11.97 | 0.4950 |

ID Balancing achieves the lowest LM loss and lowest MTP MaxVio, while Quantile Balancing gives the strongest backbone balance in this stress test.

### Downstream Performance (Table 2)

On nine benchmarks (MMLU, MMLU-Pro, SuperGPQA, MATH, GSM8K, BBH, MMMLU, EvalPlus, MultiPL-E):

- **Top-8-of-256 (3.0B/24.8B, 560B tokens)**: ID Balancing achieves Avg. 55.81 (highest among all methods)
- **Top-10-of-768 (3.2B/69.9B, 560B tokens)**: ID Balancing achieves Avg. 58.66, improving over Auxiliary loss (58.12) and scoring higher on five of nine benchmarks

### Continued Pretraining

With reduced gains ($K_i = K_d = 6 \times 10^{-6}$), ID Balancing maintains:
- LM loss within 0.001 of the frozen-bias baseline during the second half of continued pretraining
- Mean MaxVio settles near 1 (Top-8-of-256) or falls from ~6 to 1.4 (Top-10-of-768)
- Frozen baseline shows severe underload with MinVio reaching ~0.95 for Top-10-of-768

### Inference-Time Expert Utilization

Under Top-10-of-768 routing, ID Balancing reduces the average inactive-expert ratio from 8.4% (Auxiliary loss) to 6.4%, and the first-layer ratio from 9.3% to 2.6%.

### Ablation Results

**Integral gain sweep** ($K_d = 0$, Table 4):
- $K_i = 6 \times 10^{-3}$: Worst backbone MaxVio = 14.96 (lowest)
- $K_i = 9 \times 10^{-3}$: Lower training-average but worse MTP balance (Worst MTP MaxVio = 35.56 vs 15.18)

**Derivative gain sweep** ($K_i = 6 \times 10^{-3}$, Table 5):
- $K_d = 6 \times 10^{-3}$: Reduces mean backbone MaxVio from 2.0020 to 1.9330 over 0–1k steps; reduces cumulative excess overload by 57.9% over 1–5k steps relative to integral-only
- Larger gains trade off increased MTP overload

## Theoretical and Practical Implications

### Theoretical Contributions

1. **Unified PID framework**: Provides a principled lens for understanding auxiliary-loss-free MoE balancing methods, connecting their update rules to classical control theory concepts. This reveals why existing methods have complementary weaknesses: fixed-step integral control (DeepSeek) cannot scale corrections to error magnitude, while proportional control (Quantile Balancing) lacks error accumulation memory.

2. **Design principles**: The analysis motivates three design choices:
   - Corrections should scale with load error magnitude
   - Extra correction should activate only when imbalance worsens
   - Bias re-centering maintains routing invariance while preventing drift

### Practical Implications

1. **Scaling sparser MoE models**: ID Balancing's advantages grow with sparsity, making it a promising solution for scaling to larger, sparser MoE architectures where load imbalance becomes increasingly severe.

2. **Training stability**: Lower, smoother gradient norms and reduced transient overload suggest improved training dynamics, particularly important at higher learning rates.

3. **Continued pretraining**: Reduced gains enable active load correction during domain adaptation without disrupting LM loss, addressing a limitation of frozen-bias approaches.

4. **Inference efficiency**: Fewer inactive experts on evaluation data means better utilization of model capacity at inference time.

## Conclusion

The paper proposes **ID Balancing**, an Integral–Derivative controller for stable training of highly sparse MoE models. By unifying existing auxiliary-loss-free methods under a generalized PID framework, the authors identify that DeepSeek's loss-free method acts as fixed-step integral control while Quantile Balancing functions as generalized proportional control. ID Balancing combines:
- A **magnitude-aware integral term** that scales corrections with load error
- A **worsening-gated derivative term** that activates only when imbalance deteriorates
- **Zero-mean bias centering** to prevent common drift

The method requires only $O(E)$ token-count feedback per layer and adds no balancing gradient to the language-modeling objective. Experiments across routing sparsities (Top-10/5/3 of 768), model sizes (18.9B to 69.9B), and learning-rate conditions demonstrate effective load control with competitive language-modeling and downstream performance.

**Future directions** include:
- Extending evaluation to other architectures and routing methods
- Further theoretical analysis of the accumulated worsening-gated derivative update
- Isolating the effect of model size alone in scaling experiments (current comparisons vary multiple configuration factors simultaneously)

---

_Markdown view of https://picx.dev/p/DiQXOy, served by PicX — AI-generated visual whiteboard summaries of research papers._
