# Exact Quantile Balancing and Load-Error Injection for Mixture-of-Experts

> Exact Quantile Balancing and Load-Error Injection jointly improve global and local MoE load balance, cutting MaxVio up to 19.4% and 25% respectively at comparable model quality.

- **Source:** [arXiv](https://arxiv.org/abs/2609.28053)
- **Published:** 2026-10-03
- **Permalink:** https://picx.dev/p/9orG7c
- **Whiteboard:** https://picx.dev/p/9orG7c/image

## Summary

## Summary (Overview)

- **Exact Quantile Balancing (EQB)**: A novel distributed method that computes exact global-batch BF16 quantiles for MoE load balancing using two-pass radix selection with only two 256-bin all-reduces per layer, achieving token-count-independent communication.
- **Load-Error Injection (LEI)**: A new gradient-based load balancing technique that injects local load errors directly into router-score gradients, avoiding the coupling and soft-surrogate issues of the GShard auxiliary loss.
- **Complementary control**: EQB addresses global balance (across an optimizer step) while LEI addresses local balance (within expert-parallel microbatches), providing complementary solutions to the two distinct load balancing challenges in MoE training.
- **Empirical gains**: On 7.5B-parameter MoEs trained up to 500B tokens, EQB reduces Global MaxVio by 19.4% over rank-averaged QB, while LEI improves Local MaxVio from 4.71 to 3.52 compared to GShard loss at comparable quality.
- **Stability mechanism**: A bounded residual transformation ($\rho \tanh(\xi)/\xi$) preserves LEI's correction direction while preventing attention-logit instability.

---

## Introduction and Theoretical Foundation

### Background

Sparse Mixture-of-Experts (MoE) models increase parameter capacity without proportionally increasing per-token compute by routing each token to only a few experts [Shazeer et al., 2017, Fedus et al., 2022]. Realizing this benefit requires balancing expert load at **two scales**:

1. **Global balance**: Measured over all tokens in an optimizer step; prevents routing from persistently concentrating on a small subset of experts.
2. **Local balance**: Within an expert-parallel (EP) microbatch; reduces load skew, improving dispatch and expert-compute efficiency.

These objectives are distinct: imbalances across local microbatches can cancel when aggregated over an optimizer step, so global balance does not imply local balance.

### MoE Routing Formulation

For each token representation $x_t \in \mathbb{R}^d$, an MoE layer contains $E$ routed experts and selects $K$ experts per token. The router produces logits $z_t = W_r x_t \in \mathbb{R}^E$ and selects $I_t = \text{Top-}K(z_t + b)$, where $b \in \mathbb{R}^E$ is the expert bias used for load balancing [Wang et al., 2024]. The output is:

$$y_t = \sum_{e \in I_t} s_{t,e} \text{FFN}_e(x_t)$$

where $s_{t,e} = \text{sigmoid}(z_{t,e})$ are unbiased sigmoid scores. The load of expert $e$ for a token batch $B$ is:

$$f_e(B) = \sum_{t \in B} \mathbb{1}[e \in I_t]$$

with uniform target $\bar{f}(B) = K|B|/E$.

### Quantile Balancing (QB)

Quantile Balancing [Su, 2026] computes the expert bias directly from the distribution of routing margins. For optimizer step $q$ and token $t$, let $\tau_t$ denote the $(K+1)$-th largest entry of $z_t + b^{(q)}$. The QB update is:

$$\hat{b}^{(q+1)}_e = \text{quantile}_{K/E}\{\tau_t - z_{t,e} : t \in B_g\}, \quad b^{(q+1)} = \hat{b}^{(q+1)} - \text{mean}(\hat{b}^{(q+1)})\mathbb{1}$$

The centering removes a common offset that does not affect top-$K$ selection.

### Limitations of Existing Approaches

- **Rank-averaged QB** [Dial, 2026]: Computes a quantile on each shard and averages across ranks. Since quantiles do not commute with averaging, this is biased and partition-dependent.
- **Histogram-based QB** [Kimi Team, 2026]: Aggregates histogram counts across ranks, yielding partition-invariant estimates but limited by histogram resolution.
- **GShard auxiliary loss** [Lepikhin et al., 2021]: Uses normalized routing probabilities, coupling corrections across experts and using a soft surrogate that can change without changing hard top-$K$ assignments.

---

## Methodology

### Exact Quantile Balancing (EQB)

EQB adapts **radix selection** [Alabi et al., 2012, Li et al., 2024] to find the exact global-batch empirical quantile. The approach exploits the 16-bit BF16 representation with an 8-bit radix, requiring exactly two passes:

**Coarse pass**: Each rank forms 256-bin counts of high bytes locally; an all-reduce locates the target high byte $u_e$ globally and its rank $\rho_e$ within that range.

**Fine pass**: Each rank counts low bytes only for margins with high byte $u_e$; cumulative counts identify the low byte $v_e$ at rank $\rho_e$.

Since the encoding preserves BF16 order, $(u_e, v_e)$ recovers the exact empirical quantile. The result is **invariant to how tokens are partitioned** across data-parallel ranks because both passes aggregate global counts.

**Communication cost**: $2E \cdot 256$ int32 counts per layer, independent of the number of tokens. For $E = 256$, this is 0.5 MiB per layer, or 24 MiB per optimizer step for 48 MoE layers — **128× smaller** than all-reducing a full $2^{16}$-bin BF16 histogram.

### Load-Error Injection (LEI)

The natural load-balancing objective is:

$$L_{bal} = \frac{1}{2}\|F - Q\|_2^2$$

where $F_e = f_e(B_l)/(K|B_l|)$ is the hard-load fraction and $Q_e = 1/E$ is the uniform target. However, $F$ is determined by discrete top-$K$ assignment and has zero gradient almost everywhere.

Using a straight-through estimator (STE) with differentiable surrogate $P(s)$:

$$\tilde{F} = P(s) + \text{sg}(F - P(s)), \quad L^{STE}_{bal} = \frac{1}{2}\|\tilde{F} - Q\|_2^2$$

The backward pass yields:

$$\nabla_s L^{STE}_{bal} = J_{P(s)}^\top (F - Q)$$

**LEI** uses the mean unnormalized router score as the STE surrogate:

$$P^L_e = \frac{1}{|B_l|}\sum_{t \in B_l} s_{t,e}$$

Its Jacobian is constant and diagonal:

$$\frac{\partial P^L_e}{\partial s_{t,e'}} = \frac{1}{|B_l|}\mathbb{1}[e = e']$$

Defining the relative load error $\rho_e = f_e(B_l)/\bar{f}(B_l) - 1 = E(F_e - Q_e)$, LEI applies the simple gradient correction:

$$\text{grad}_{s_{t,e}} = \text{grad}_{s_{t,e}} + \eta \rho_e, \quad t \in B_l$$

where $\eta > 0$ controls correction strength. Overloaded experts receive positive score gradients; underloaded experts receive negative ones.

### Stabilization

To prevent attention-logit instability from unbounded auxiliary gradients, LEI bounds the injected residual while preserving its STE direction:

$$\tilde{\rho} = \rho \frac{\tanh(\xi)}{\xi}, \quad \xi = \frac{\|\rho\|_\infty}{c}$$

with $c > 0$ and $\tanh(\xi)/\xi := 1$ at $\xi = 0$. This preserves direction and zero-centering while bounding the injected gradient.

---

## Empirical Validation / Results

### Experimental Setup

- **Model**: 7.5B-parameter decoder-only MoE with 0.5B active parameters per token
- **Configuration**: $E = 256$ routed experts, $K = 6$
- **Training**: Six ablations for 100B tokens; EQB vs. normalized LEI compared at 500B tokens
- **Metric**: MaxVio$(B) = \max_e[f_e(B)/\bar{f}(B) - 1]$ [DeepSeek-AI, 2024], with Global MaxVio over an optimizer step and Local MaxVio on rank-local EP microbatches

### Table 1: 100B-Token Ablation Results

| Method | MaxVio Global ↓ | MaxVio Local ↓ | MMLU | ARC | HSwag | GSM8K | HumanEval | MBPP | Mean |
|--------|-----------------|----------------|------|-----|-------|-------|-----------|------|------|
| Rank-avg. QB | 0.92 | 5.96 | 0.9007 | 0.7508 | 0.7718 | 0.5260 | 0.5205 | 0.5993 | 0.6782 |
| **EQB** | **0.74** | **5.38** | 0.8650 | 0.6876 | 0.7703 | 0.5258 | 0.5154 | 0.5777 | 0.6570 |
| EQB + GShard loss | 0.71 | 4.71 | 0.8932 | 0.7295 | 0.7737 | 0.5400 | 0.5129 | 0.6625 | 0.6853 |
| **EQB + LEI** | **0.60** | **3.49** | 0.8897 | 0.7348 | 0.7729 | 0.5409 | 0.4989 | 0.6177 | 0.6758 |
| EQB + norm. GShard loss | 0.74 | 4.30 | 0.8853 | 0.7373 | 0.7730 | 0.5333 | 0.5031 | 0.6228 | 0.6758 |
| EQB + norm. LEI | 0.60 | 3.52 | 0.8892 | 0.7198 | 0.7729 | 0.5341 | 0.5167 | 0.5992 | 0.6720 |

### Table 2: 500B-Token EQB vs. Normalized LEI

| Method | MaxVio Global ↓ | MaxVio Local ↓ | MMLU | ARC | HSwag | TriviaQA | Avg. |
|--------|-----------------|----------------|------|-----|-------|----------|------|
| EQB | 0.44 | 5.14 | 41.36 | 50.05 | 63.53 | 38.06 | 48.25 |
| **EQB + norm. LEI** | 0.63 | **4.15** | 41.38 | **51.49** | 63.53 | 37.83 | **48.56** |

### Key Findings

1. **EQB improves on rank-averaged QB**: Replacing rank-averaged quantiles with the exact global-batch statistic reduces Global MaxVio from 0.92 to 0.74 (19.4%) and Local MaxVio from 5.96 to 5.38 (9.8%), while improving BPB on all six reported benchmarks.

2. **LEI improves local balance**: At 100B tokens, normalized LEI reduces Global/Local MaxVio to 0.60/3.52, compared with 0.71/4.71 for the GShard loss, at comparable model quality.

3. **Longer training**: At 500B tokens, LEI reduces Local MaxVio from 5.14 to 4.15 while average accuracy increases from 48.25% to 48.56%; Global MaxVio increases from 0.44 to 0.63.

---

## Theoretical and Practical Implications

### Theoretical Contributions

- **Partition-invariant exact quantiles**: EQB demonstrates that exact global-batch order statistics can be recovered in distributed settings without materializing token-level data centrally, using the BF16 representation's structure for efficient radix selection.
- **STE surrogate analysis**: The paper provides a unified view of GShard and LEI as straight-through estimators with different surrogate choices, clarifying why the diagonal, constant Jacobian of LEI avoids the coupling issues of probability-normalized surrogates.
- **Complementary control scales**: The work formalizes the distinction between global and local load balance, showing that token-independent biases and gradient-based corrections address fundamentally different imbalance sources.

### Practical Implications

- **Communication efficiency**: EQB's communication cost is independent of token count and 128× smaller than full-histogram approaches, making it practical for large-scale training.
- **Stability**: The bounded residual transformation ($\rho \tanh(\xi)/\xi$) provides a practical recipe for stable auxiliary gradient injection without hyperparameter-sensitive normalization.
- **Deployment**: The methods are drop-in replacements for existing QB and GShard loss components, requiring no changes to forward routing or mixture weights.

---

## Conclusion

EQB and LEI address complementary scales of MoE load balancing:

- **EQB** recovers the exact global-batch quantile with token-count-independent communication, improving on rank-averaged QB in both balance metrics and downstream performance.
- **LEI** injects local load residuals directly into router-score gradients, improving balance over the GShard loss at comparable quality.
- Bounding the injected residual preserves improvement without attention-logit instability.

### Limitations

- Results use one model scale and seed, without estimating variance.
- GShard and LEI use separately tuned coefficients because the same coefficient produces very different router-gradient magnitudes.
- Local MaxVio is a proxy for dispatch cost rather than measured EP throughput.

### Future Directions

The authors' framework suggests several promising extensions: applying EQB and LEI to larger model scales with variance estimation, developing adaptive coefficient schemes that handle the different gradient magnitudes automatically, and validating local balance improvements against direct EP throughput measurements.

---

_Markdown view of https://picx.dev/p/9orG7c, served by PicX — AI-generated visual whiteboard summaries of research papers._
