Full text not available for this paper
Summary (Overview)
- Exact Quantile Balancing (EQB): A novel distributed method that computes exact global-batch BF16 quantiles for MoE load balancing using two-pass radix selection with only two 256-bin all-reduces per layer, achieving token-count-independent communication.
- Load-Error Injection (LEI): A new gradient-based load balancing technique that injects local load errors directly into router-score gradients, avoiding the coupling and soft-surrogate issues of the GShard auxiliary loss.
- Complementary control: EQB addresses global balance (across an optimizer step) while LEI addresses local balance (within expert-parallel microbatches), providing complementary solutions to the two distinct load balancing challenges in MoE training.
- Empirical gains: On 7.5B-parameter MoEs trained up to 500B tokens, EQB reduces Global MaxVio by 19.4% over rank-averaged QB, while LEI improves Local MaxVio from 4.71 to 3.52 compared to GShard loss at comparable quality.
- Stability mechanism: A bounded residual transformation () preserves LEI's correction direction while preventing attention-logit instability.
Introduction and Theoretical Foundation
Background
Sparse Mixture-of-Experts (MoE) models increase parameter capacity without proportionally increasing per-token compute by routing each token to only a few experts [Shazeer et al., 2017, Fedus et al., 2022]. Realizing this benefit requires balancing expert load at two scales:
- Global balance: Measured over all tokens in an optimizer step; prevents routing from persistently concentrating on a small subset of experts.
- Local balance: Within an expert-parallel (EP) microbatch; reduces load skew, improving dispatch and expert-compute efficiency.
These objectives are distinct: imbalances across local microbatches can cancel when aggregated over an optimizer step, so global balance does not imply local balance.
MoE Routing Formulation
For each token representation , an MoE layer contains routed experts and selects experts per token. The router produces logits and selects , where is the expert bias used for load balancing [Wang et al., 2024]. The output is:
where are unbiased sigmoid scores. The load of expert for a token batch is:
with uniform target .
Quantile Balancing (QB)
Quantile Balancing [Su, 2026] computes the expert bias directly from the distribution of routing margins. For optimizer step and token , let denote the -th largest entry of . The QB update is:
The centering removes a common offset that does not affect top- selection.
Limitations of Existing Approaches
- Rank-averaged QB [Dial, 2026]: Computes a quantile on each shard and averages across ranks. Since quantiles do not commute with averaging, this is biased and partition-dependent.
- Histogram-based QB [Kimi Team, 2026]: Aggregates histogram counts across ranks, yielding partition-invariant estimates but limited by histogram resolution.
- GShard auxiliary loss [Lepikhin et al., 2021]: Uses normalized routing probabilities, coupling corrections across experts and using a soft surrogate that can change without changing hard top- assignments.
Methodology
Exact Quantile Balancing (EQB)
EQB adapts radix selection [Alabi et al., 2012, Li et al., 2024] to find the exact global-batch empirical quantile. The approach exploits the 16-bit BF16 representation with an 8-bit radix, requiring exactly two passes:
Coarse pass: Each rank forms 256-bin counts of high bytes locally; an all-reduce locates the target high byte globally and its rank within that range.
Fine pass: Each rank counts low bytes only for margins with high byte ; cumulative counts identify the low byte at rank .
Since the encoding preserves BF16 order, recovers the exact empirical quantile. The result is invariant to how tokens are partitioned across data-parallel ranks because both passes aggregate global counts.
Communication cost: int32 counts per layer, independent of the number of tokens. For , this is 0.5 MiB per layer, or 24 MiB per optimizer step for 48 MoE layers — 128× smaller than all-reducing a full -bin BF16 histogram.
Load-Error Injection (LEI)
The natural load-balancing objective is:
where is the hard-load fraction and is the uniform target. However, is determined by discrete top- assignment and has zero gradient almost everywhere.
Using a straight-through estimator (STE) with differentiable surrogate :
The backward pass yields:
LEI uses the mean unnormalized router score as the STE surrogate:
Its Jacobian is constant and diagonal:
Defining the relative load error , LEI applies the simple gradient correction:
where controls correction strength. Overloaded experts receive positive score gradients; underloaded experts receive negative ones.
Stabilization
To prevent attention-logit instability from unbounded auxiliary gradients, LEI bounds the injected residual while preserving its STE direction:
with and at . This preserves direction and zero-centering while bounding the injected gradient.
Empirical Validation / Results
Experimental Setup
- Model: 7.5B-parameter decoder-only MoE with 0.5B active parameters per token
- Configuration: routed experts,
- Training: Six ablations for 100B tokens; EQB vs. normalized LEI compared at 500B tokens
- Metric: MaxVio [DeepSeek-AI, 2024], with Global MaxVio over an optimizer step and Local MaxVio on rank-local EP microbatches
Table 1: 100B-Token Ablation Results
| Method | MaxVio Global ↓ | MaxVio Local ↓ | MMLU | ARC | HSwag | GSM8K | HumanEval | MBPP | Mean |
|---|---|---|---|---|---|---|---|---|---|
| Rank-avg. QB | 0.92 | 5.96 | 0.9007 | 0.7508 | 0.7718 | 0.5260 | 0.5205 | 0.5993 | 0.6782 |
| EQB | 0.74 | 5.38 | 0.8650 | 0.6876 | 0.7703 | 0.5258 | 0.5154 | 0.5777 | 0.6570 |
| EQB + GShard loss | 0.71 | 4.71 | 0.8932 | 0.7295 | 0.7737 | 0.5400 | 0.5129 | 0.6625 | 0.6853 |
| EQB + LEI | 0.60 | 3.49 | 0.8897 | 0.7348 | 0.7729 | 0.5409 | 0.4989 | 0.6177 | 0.6758 |
| EQB + norm. GShard loss | 0.74 | 4.30 | 0.8853 | 0.7373 | 0.7730 | 0.5333 | 0.5031 | 0.6228 | 0.6758 |
| EQB + norm. LEI | 0.60 | 3.52 | 0.8892 | 0.7198 | 0.7729 | 0.5341 | 0.5167 | 0.5992 | 0.6720 |
Table 2: 500B-Token EQB vs. Normalized LEI
| Method | MaxVio Global ↓ | MaxVio Local ↓ | MMLU | ARC | HSwag | TriviaQA | Avg. |
|---|---|---|---|---|---|---|---|
| EQB | 0.44 | 5.14 | 41.36 | 50.05 | 63.53 | 38.06 | 48.25 |
| EQB + norm. LEI | 0.63 | 4.15 | 41.38 | 51.49 | 63.53 | 37.83 | 48.56 |
Key Findings
-
EQB improves on rank-averaged QB: Replacing rank-averaged quantiles with the exact global-batch statistic reduces Global MaxVio from 0.92 to 0.74 (19.4%) and Local MaxVio from 5.96 to 5.38 (9.8%), while improving BPB on all six reported benchmarks.
-
LEI improves local balance: At 100B tokens, normalized LEI reduces Global/Local MaxVio to 0.60/3.52, compared with 0.71/4.71 for the GShard loss, at comparable model quality.
-
Longer training: At 500B tokens, LEI reduces Local MaxVio from 5.14 to 4.15 while average accuracy increases from 48.25% to 48.56%; Global MaxVio increases from 0.44 to 0.63.
Theoretical and Practical Implications
Theoretical Contributions
- Partition-invariant exact quantiles: EQB demonstrates that exact global-batch order statistics can be recovered in distributed settings without materializing token-level data centrally, using the BF16 representation's structure for efficient radix selection.
- STE surrogate analysis: The paper provides a unified view of GShard and LEI as straight-through estimators with different surrogate choices, clarifying why the diagonal, constant Jacobian of LEI avoids the coupling issues of probability-normalized surrogates.
- Complementary control scales: The work formalizes the distinction between global and local load balance, showing that token-independent biases and gradient-based corrections address fundamentally different imbalance sources.
Practical Implications
- Communication efficiency: EQB's communication cost is independent of token count and 128× smaller than full-histogram approaches, making it practical for large-scale training.
- Stability: The bounded residual transformation () provides a practical recipe for stable auxiliary gradient injection without hyperparameter-sensitive normalization.
- Deployment: The methods are drop-in replacements for existing QB and GShard loss components, requiring no changes to forward routing or mixture weights.
Conclusion
EQB and LEI address complementary scales of MoE load balancing:
- EQB recovers the exact global-batch quantile with token-count-independent communication, improving on rank-averaged QB in both balance metrics and downstream performance.
- LEI injects local load residuals directly into router-score gradients, improving balance over the GShard loss at comparable quality.
- Bounding the injected residual preserves improvement without attention-logit instability.
Limitations
- Results use one model scale and seed, without estimating variance.
- GShard and LEI use separately tuned coefficients because the same coefficient produces very different router-gradient magnitudes.
- Local MaxVio is a proxy for dispatch cost rather than measured EP throughput.
Future Directions
The authors' framework suggests several promising extensions: applying EQB and LEI to larger model scales with variance estimation, developing adaptive coefficient schemes that handle the different gradient magnitudes automatically, and validating local balance improvements against direct EP throughput measurements.
Related papers
- SOLAR: A State-Driven Online Learning Rate Scheduler for LLM Pretraining
SOLAR uses reinforcement learning to learn bounded residual corrections to a base learning-rate schedule, improving LLM pretraining perplexity across dense and MoE models up to 3B parameters.
- How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining
Encoder-free multimodal LLMs match encoder-based performance at ~10^22 FLOPs, shifting compute-optimal allocation toward larger decoders and enabling viable encoder-free pretraining.
- Just Let Linear States Forget the Distant Past: Prefix Caching via Suffix Replay for Hybrid LLMs
SuffixReplay enables fine-grained prefix caching in hybrid LLMs by reconstructing linear-attention states from sparse anchors, cutting storage by 2x and TTFT by up to 70%.