Full text not available for this paper

Summary (Overview)

  • Unified PID framework: The paper proposes a generalized Proportional–Integral–Derivative (PID) control view of auxiliary-loss-free MoE load balancing, showing that DeepSeek's loss-free method acts as a fixed-step integral controller while Kimi K3's Quantile Balancing functions as a generalized proportional controller.
  • ID Balancing method: Introduces an Integral–Derivative controller combining a magnitude-aware integral term with a worsening-gated derivative term, requiring only O(E)O(E) token-count feedback per layer with no balancing gradient added to the training objective.
  • Superior load control: Reduces worst-case backbone MaxVio by over 50% and training-average backbone MinVio by over 12% relative to the best baselines in the Top-3-of-768 setting, with advantages growing as sparsity increases.
  • Scale robustness: Under fixed Top-10-of-768 routing, worst-case backbone MaxVio remains nearly unchanged as total parameters increase from 18.9B to 69.9B (approximately 89.6% lower than the auxiliary-loss baseline).
  • Competitive performance: Maintains competitive language-modeling loss and downstream performance across nine benchmarks while achieving effective load control.

Introduction and Theoretical Foundation

Background and Motivation

Mixture-of-Experts (MoE) enables efficient scaling of large language models through sparse expert activation. For each token, a router selects only KK experts from a pool of EE. With KK fixed, adding experts increases total capacity without increasing per-token computation. However, as the active fraction K/EK/E decreases, expert load imbalance becomes more severe:

  • Overloaded experts slow down MoE computation
  • Underloaded experts receive insufficient training, leaving model capacity unused

Control-Theoretic Formulation

The paper formulates MoE load balancing as a control problem. For an MoE layer with EE routed experts and a Top-KK router:

  • Router logits: z=Wrxz = W_r x where Wr∈RE×dW_r \in \mathbb{R}^{E \times d} is the router weight matrix
  • Expert scores: si=σ(zi)s_i = \sigma(z_i) via activation function σ\sigma
  • Expert selection with bias: T(x)=TopKi∈[E](si+bi)T(x) = \text{TopK}_{i \in [E]}(s_i + b_i)
  • Normalized load error: ei(t)=nˉ(t)−ni(t)nˉ(t)e_i^{(t)} = \frac{\bar{n}^{(t)} - n_i^{(t)}}{\bar{n}^{(t)}} where nˉ(t)=1E∑i=1Eni(t)\bar{n}^{(t)} = \frac{1}{E}\sum_{i=1}^{E} n_i^{(t)}

The standard PID controller is given by:

ut=Kpet+Ki∑τ≤teτ+Kd(et−et−1)u_t = K_p e_t + K_i \sum_{\tau \leq t} e_\tau + K_d(e_t - e_{t-1})

where Kp,Ki,Kd≥0K_p, K_i, K_d \geq 0 are the gains of the proportional, integral, and derivative branches.

Existing Methods Through a PID Lens

Under the generalized PID view (where temporal roles define branches, allowing nonlinear transformations):

DeepSeek's loss-free method (fixed-step integral control):

b(t+1)=b(t)+η⋅sign(e(t))=η∑τ≤tsign(e(τ))=I(t)b^{(t+1)} = b^{(t)} + \eta \cdot \text{sign}(e^{(t)}) = \eta \sum_{\tau \leq t} \text{sign}(e^{(\tau)}) = I^{(t)}

The sign operation discards error magnitude—every nonzero error receives the same correction η\eta.

Quantile Balancing (generalized proportional control):

b(t+1)=Q(S(t),b(t))=P(t)b^{(t+1)} = Q(S^{(t)}, b^{(t)}) = P^{(t)}

where Qj(S(t),b(t))=−quantile1−K/E({Si,j(t)−αi(t)}i=1m)Q_j(S^{(t)}, b^{(t)}) = -\text{quantile}_{1-K/E}\left(\{S_{i,j}^{(t)} - \alpha_i^{(t)}\}_{i=1}^{m}\right)

This directly responds to batch-dependent targets without explicitly accumulating past errors.

Methodology

ID Balancing Algorithm

The proposed ID Balancing combines two key components:

1. Magnitude-aware integral term:

bi(t+1)=bi(t)+Kiei(t)b^{(t+1)}_i = b^{(t)}_i + K_i e^{(t)}_i

The correction grows with ∣ei(t)∣|e^{(t)}_i| and shrinks near balance. Unlike the sign-based update, this preserves zero-mean bias when initialized at b(0)=0b^{(0)} = 0 since ∑iei(t)=0\sum_i e^{(t)}_i = 0.

2. Worsening-gated derivative term:

Δei(t)=ei(t)−ei(t−1)\Delta e^{(t)}_i = e^{(t)}_i - e^{(t-1)}_i gi(t)=1[ei(t−1)⋅Δei(t)>0]g^{(t)}_i = \mathbb{1}[e^{(t-1)}_i \cdot \Delta e^{(t)}_i > 0]

The gate opens only when the imbalance worsens (error and its change have the same sign), providing extra correction only during deterioration.

Complete update with re-centering:

b~i(t+1)=bi(t)+Kiei(t)+Kdgi(t)Δei(t)\tilde{b}^{(t+1)}_i = b^{(t)}_i + K_i e^{(t)}_i + K_d g^{(t)}_i \Delta e^{(t)}_i bi(t+1)=b~i(t+1)−1E∑j=1Eb~j(t+1)b^{(t+1)}_i = \tilde{b}^{(t+1)}_i - \frac{1}{E}\sum_{j=1}^{E} \tilde{b}^{(t+1)}_j

The zero-mean centering removes common bias drift without changing routing decisions (adding the same constant to all biases leaves the ranking unchanged).

Algorithm Pseudocode

Algorithm 1: ID Balancing bias update, one MoE layer, per training iteration
Require: expert count E, active experts K, integral gain K_i, derivative gain K_d
1: b_i ← 0, e_prev_i ← 0 for all i ∈ [E]
2: for each training iteration do
3:   route the batch by T(x) = TopK_{i∈[E]}(s_i + b_i)
4:   n_i ← number of tokens assigned to expert i
5:   n̄ ← (1/E) Σ_{i=1}^{E} n_i
6:   for each expert i ∈ [E] do
7:     e_i ← (n̄ - n_i) / n̄
8:     Δe_i ← e_i - e_prev_i
9:     g_i ← 1[e_prev_i · Δe_i > 0]
10:    b_i ← b_i + K_i e_i + K_d g_i Δe_i
11:    e_prev_i ← e_i
12:  end for
13:  b_i ← b_i - (1/E) Σ_{j=1}^{E} b_j for all i ∈ [E]
14: end for

Experimental Setup

  • Architecture: Qwen-3.8-Next with 20 Transformer layers, one multi-token prediction (MTP) module, 768 routed experts per MoE layer
  • Model sizes: 18.9B total parameters (0.87–1.03B active), 24.8B, and 69.9B total parameters
  • Routing settings: Top-10, Top-5, and Top-3 over 768 experts
  • Training: 120B tokens (main comparison), 560B tokens (downstream evaluation), global batch size 1024
  • Learning rate: Peak 2.54 × 10⁻³ with cosine decay to 3 × 10⁻⁵; stress test at constant 5.86 × 10⁻³ (2.3× standard peak)
  • Default gains: Ki=Kd=6×10−3K_i = K_d = 6 \times 10^{-3} (reduced to 6×10−66 \times 10^{-6} for continued pretraining)

Evaluation Metrics

MaxVio(t)=max⁡ini(t)−nˉ(t)nˉ(t),MinVio(t)=nˉ(t)−min⁡ini(t)nˉ(t)\text{MaxVio}^{(t)} = \max_i \frac{n_i^{(t)} - \bar{n}^{(t)}}{\bar{n}^{(t)}}, \quad \text{MinVio}^{(t)} = \frac{\bar{n}^{(t)} - \min_i n_i^{(t)}}{\bar{n}^{(t)}}
  • MaxVio: largest relative overload (bounded by E/K−1E/K - 1)
  • MinVio: largest relative deficit (in [0, 1], reaching 1 when an expert receives no tokens)

Empirical Validation / Results

Main Results (Table 1): Load Control Across Routing Sparsities

MethodAct./TotalLM lossBackbone MaxVio (Worst)Backbone MinVio (Avg.)
Top-3 of 768 experts0.87B/18.9B
Auxiliary loss1.752770.760.6232
DeepSeek loss-free1.7426211.190.7441
Quantile Balancing1.743231.110.6373
ID Balancing1.742615.330.5480
Top-5 of 768 experts0.91B/18.9B
Auxiliary loss1.737850.480.5836
DeepSeek loss-free1.7280115.820.6725
Quantile Balancing1.730321.160.5232
ID Balancing1.729910.330.4601
Top-10 of 768 experts1.03B/18.9B
Auxiliary loss1.718016.420.5446
DeepSeek loss-free1.712169.070.5571
Quantile Balancing1.713310.130.3914
ID Balancing1.71188.360.3800

Key findings: ID Balancing achieves the lowest Worst and Last1k MaxVio, as well as the lowest Last1k and Avg. MinVio across all three routing settings, with LM loss gaps under 0.0023 among bias-based methods.

Scaling Results (Figure 1)

  • Sparsity scaling: As routing becomes sparser (Top-10 → Top-3 of 768), ID Balancing's worst-case MaxVio grows from 8.36 to 15.33, while DeepSeek's grows from 69.07 to 211.19
  • Model size scaling: Under fixed Top-10-of-768 routing, ID Balancing's worst-case MaxVio remains nearly unchanged (8.36 → 7.79) as total parameters increase from 18.9B to 69.9B, approximately 89.6% lower than the auxiliary-loss baseline

Higher Learning Rate Stress Test (Table 3)

At constant learning rate 5.86 × 10⁻³ (2.3× standard peak):

MethodLM lossBackbone MaxVio (Worst)Backbone MinVio (Avg.)
Auxiliary loss2.009920.890.8458
DeepSeek loss-free2.003172.540.8269
Quantile Balancing2.00092.960.3938
ID Balancing2.000711.970.4950

ID Balancing achieves the lowest LM loss and lowest MTP MaxVio, while Quantile Balancing gives the strongest backbone balance in this stress test.

Downstream Performance (Table 2)

On nine benchmarks (MMLU, MMLU-Pro, SuperGPQA, MATH, GSM8K, BBH, MMMLU, EvalPlus, MultiPL-E):

  • Top-8-of-256 (3.0B/24.8B, 560B tokens): ID Balancing achieves Avg. 55.81 (highest among all methods)
  • Top-10-of-768 (3.2B/69.9B, 560B tokens): ID Balancing achieves Avg. 58.66, improving over Auxiliary loss (58.12) and scoring higher on five of nine benchmarks

Continued Pretraining

With reduced gains (Ki=Kd=6×10−6K_i = K_d = 6 \times 10^{-6}), ID Balancing maintains:

  • LM loss within 0.001 of the frozen-bias baseline during the second half of continued pretraining
  • Mean MaxVio settles near 1 (Top-8-of-256) or falls from ~6 to 1.4 (Top-10-of-768)
  • Frozen baseline shows severe underload with MinVio reaching ~0.95 for Top-10-of-768

Inference-Time Expert Utilization

Under Top-10-of-768 routing, ID Balancing reduces the average inactive-expert ratio from 8.4% (Auxiliary loss) to 6.4%, and the first-layer ratio from 9.3% to 2.6%.

Ablation Results

Integral gain sweep (Kd=0K_d = 0, Table 4):

  • Ki=6×10−3K_i = 6 \times 10^{-3}: Worst backbone MaxVio = 14.96 (lowest)
  • Ki=9×10−3K_i = 9 \times 10^{-3}: Lower training-average but worse MTP balance (Worst MTP MaxVio = 35.56 vs 15.18)

Derivative gain sweep (Ki=6×10−3K_i = 6 \times 10^{-3}, Table 5):

  • Kd=6×10−3K_d = 6 \times 10^{-3}: Reduces mean backbone MaxVio from 2.0020 to 1.9330 over 0–1k steps; reduces cumulative excess overload by 57.9% over 1–5k steps relative to integral-only
  • Larger gains trade off increased MTP overload

Theoretical and Practical Implications

Theoretical Contributions

  1. Unified PID framework: Provides a principled lens for understanding auxiliary-loss-free MoE balancing methods, connecting their update rules to classical control theory concepts. This reveals why existing methods have complementary weaknesses: fixed-step integral control (DeepSeek) cannot scale corrections to error magnitude, while proportional control (Quantile Balancing) lacks error accumulation memory.

  2. Design principles: The analysis motivates three design choices:

    • Corrections should scale with load error magnitude
    • Extra correction should activate only when imbalance worsens
    • Bias re-centering maintains routing invariance while preventing drift

Practical Implications

  1. Scaling sparser MoE models: ID Balancing's advantages grow with sparsity, making it a promising solution for scaling to larger, sparser MoE architectures where load imbalance becomes increasingly severe.

  2. Training stability: Lower, smoother gradient norms and reduced transient overload suggest improved training dynamics, particularly important at higher learning rates.

  3. Continued pretraining: Reduced gains enable active load correction during domain adaptation without disrupting LM loss, addressing a limitation of frozen-bias approaches.

  4. Inference efficiency: Fewer inactive experts on evaluation data means better utilization of model capacity at inference time.

Conclusion

The paper proposes ID Balancing, an Integral–Derivative controller for stable training of highly sparse MoE models. By unifying existing auxiliary-loss-free methods under a generalized PID framework, the authors identify that DeepSeek's loss-free method acts as fixed-step integral control while Quantile Balancing functions as generalized proportional control. ID Balancing combines:

  • A magnitude-aware integral term that scales corrections with load error
  • A worsening-gated derivative term that activates only when imbalance deteriorates
  • Zero-mean bias centering to prevent common drift

The method requires only O(E)O(E) token-count feedback per layer and adds no balancing gradient to the language-modeling objective. Experiments across routing sparsities (Top-10/5/3 of 768), model sizes (18.9B to 69.9B), and learning-rate conditions demonstrate effective load control with competitive language-modeling and downstream performance.

Future directions include:

  • Extending evaluation to other architectures and routing methods
  • Further theoretical analysis of the accumulated worsening-gated derivative update
  • Isolating the effect of model size alone in scaling experiments (current comparisons vary multiple configuration factors simultaneously)

Related papers