Summary (Overview)
- Core contribution: Zepp is a distributed MoE serving system that rethinks load balancing as a constraint rather than an optimization objective, directly optimizing bottleneck inter-node communication instead.
- Key insight: The paper demonstrates that "balance is not free" — operations introduced to balance one dimension (computation, communication, or memory) often introduce imbalance or overhead in another, creating a fundamental balancing dilemma.
- Three-stage optimization: Zepp progressively avoids inter-node communication (via expert replication, placement, and locality-preferred routing), reshapes it (via split and merge primitives under NIC constraints), and hides it (via communication-computation overlap scheduling).
- Dynamic adaptation: Intra-node expert swaps (rather than global remapping) rebalance GPU compute load under dynamic workloads while avoiding expensive inter-node weight transfers.
- Results: Zepp achieves up to 6.68× MoE layer speedup over existing baselines and a geometric mean speedup of 1.86× over the fastest competing baseline, with sustained gains on up to 32 nodes.
Introduction and Theoretical Foundation
Background
Mixture-of-Experts (MoE) architectures have become central to scaling large language models (LLMs), enabling models like Kimi K3 (2.8 trillion total parameters, activating only 104 billion per token) to grow without proportionally increasing per-token computation. Serving these models relies on expert parallelism (EP), which distributes experts across GPUs and routes each token to the GPUs hosting its selected experts.
The Balancing Dilemma
The paper identifies a fundamental problem: dynamic token-to-expert routing creates highly skewed and time-varying workloads across experts, GPUs, and network links. Existing systems follow a common principle — better balance leads to better performance — but the authors demonstrate this principle is flawed:
"Eliminating one form of imbalance often introduces another source of imbalance or overhead, making balance itself increasingly costly to pursue."
Specifically:
- Communication overlap (Figure 3a→3b): Hiding communication latency exposes computation imbalance as the next bottleneck
- Computation balancing (Figure 3b→3c): Expert replication and remapping introduces inter-node weight transfers that become the new bottleneck
Key Theoretical Insight
Better balance does not necessarily translate into better performance, particularly when achieving balance requires additional inter-node communication. Zepp therefore treats balance as a constraint on physical resources (GPUs and NICs) and directly minimizes the bottleneck inter-node communication.
Methodology
4.1 Expert Placement and Token Routing
Expert replication: Zepp reserves extra slots per GPU for expert replicas, greedily assigning them to experts with highest historical token demand per replica.
Hierarchical placement: Replicas are distributed round-robin across nodes first, then within nodes across GPUs, maximizing node-level expert coverage.
Locality-preferred routing: Formulated as an optimization problem:
Where is the number of tokens for expert from GPU assigned to GPU , is the demand, is the balanced reference load, is inverse bandwidth, and is a relaxation ratio controlling deviation from perfect balance.
4.2 Split and Merge for Communication Restructuring
Split: Decomposes all-to-all collectives into independently schedulable flows, then redistributes traffic across NICs under load constraints:
For homogeneous NICs with , the minimum intra-node traffic movement is:
Merge: Eliminates redundant inter-node payloads:
- Dispatch: Multiple experts selected by the same token on the same remote node share a single transfer
- Combine: Node-local pre-reduction aggregates expert outputs into partial sums before transfer
4.3 Three-Stream Scheduling
Overlapping schedule (Algorithm 1):
- Dispatch follows communication: Computation groups execute as soon as their inputs arrive
- Combine drives communication: Computation is ordered by destination groups to release complete return transfers early
Intra-node expert swaps for dynamic rebalancing:
Swaps exchange expert replicas between GPUs within the same node, preserving node-level coverage while avoiding inter-node weight transfers. Expert-weight availability is incorporated into the dependency-aware schedule.
Empirical Validation / Results
Experimental Setup
- Hardware: 4-node clusters with NVIDIA A100 (40GB/80GB) and H100 (96GB) GPUs, NVLink intra-node (100-150 GB/s), HPE Slingshot 11 NICs (25 GB/s inter-node)
- Baselines: A2AV+GEMM, FAST+GEMM, COMET, EPLB, EPIC, MoonEP, COMET+EPLB
- Workloads: Kimi-K2-Thinking (384 routed experts, top-8) and Qwen3-235B-A22B (128 routed experts, top-8)
MoE Layer Performance (Figure 9)
| Configuration | Zepp Speedup vs. A2AV+GEMM | Zepp Speedup vs. Fastest Baseline |
|---|---|---|
| 4 nodes, 16 MB (Qwen) | 4.31× | 2.18× vs. COMET+EPLB |
| 16 nodes, 1 MB (Qwen) | 6.68× | — |
| 16 nodes, 16 MB (Qwen) | 4.66× | 3.53× vs. COMET+EPLB |
| 16 nodes, 16 MB (K2) | 3.28× | 2.94× vs. COMET+EPLB |
| Geometric mean (all configs) | — | 1.86× |
Weak Scalability (Figure 10)
- At 32 nodes, 1 MB workload: 9.49× speedup on A100, 15.10× on H100
- At 64 MB workload: sustained 2.35–2.93× (A100) and 2.46–4.13× (H100) across 2–32 nodes
End-to-End Performance (Figure 11, SGLang integration)
- Prefill throughput gains: 1.32× (1 MB) to 1.92× (16 MB)
- Decode latency improvements: 1.04× (1 MB) to 1.57× (16 MB)
Ablation (Figure 12)
- Overlapping scheduling alone: 1.12× speedup over COMET
- Adding placement/routing: 1.25×
- Adding overlapped swaps: 1.28× (Prof. law dataset)
- Under rotating demand, swaps contribute 3.0% latency reduction vs. 1.9% under static demand
Theoretical and Practical Implications
Theoretical Implications
-
Paradigm shift: The paper challenges the prevailing "balance-first" paradigm in distributed MoE serving, showing that balance should be a constraint, not an objective. This has implications for how future systems reason about resource optimization in heterogeneous, dynamic environments.
-
Split vs. merge dichotomy: Unlike dense-model communication where splitting is always beneficial, MoE's sparse irregular traffic requires a bidirectional approach — split to alleviate NIC bottlenecks, merge to reduce redundant transfers and amortize latency.
-
Locality as a first-class concern: The work demonstrates that expert placement and token routing should be jointly optimized to maximize node-level locality, which is more valuable than perfect per-GPU balance.
Practical Implications
-
Scalability: Zepp's gains increase with cluster size (up to 15× at 32 nodes for small workloads), making it particularly relevant for large-scale MoE deployments.
-
Hardware efficiency: By minimizing inter-node traffic, Zepp reduces pressure on expensive scale-out networks, potentially lowering infrastructure costs.
-
Dynamic workload handling: The intra-node swap mechanism provides lightweight adaptation to changing demand patterns without the overhead of global remapping.
-
Integration feasibility: End-to-end gains in SGLang (up to 1.92× prefill throughput) demonstrate practical deployability in real serving systems.
Conclusion
Zepp rethinks distributed MoE serving by treating balance as a constraint on physical resources rather than an optimization objective. Its three-stage approach — avoiding, reshaping, and hiding inter-node communication — achieves up to 6.68× speedup over existing baselines and 1.86× geometric mean speedup over the fastest competing system.
Key contributions:
- Analysis showing how balancing one dimension shifts bottlenecks to another
- A coordinated system design combining expert replication, placement, routing, communication restructuring, and overlap scheduling
- Comprehensive evaluation across models, topologies, and workloads
Future directions suggested by the work:
- Extending the balance-as-constraint principle to other distributed serving scenarios
- Further optimization of intra-node expert swaps for more extreme demand shifts
- Exploring the interaction between MoE layer optimization and attention/other model components
- Applying split/merge transformations to other irregular communication patterns in distributed ML
The code is publicly available at github.com/andronius-yang/zepp.
Related papers
- Do Tool Calls Execute as Intended? Measuring and Repairing Intent-Execution Correspondence in LLM Agents
Tool calls are silently altered by execution-path hops in 12% of production shell invocations, and the IntAct repair protocol recovers 79.2% of resulting failures.
- More Value per Key: Asymmetric Sparse Attention for Faster LLM Decoding
SAGA decouples key and value head counts to exploit sparse attention's shifted bottleneck, achieving over 2x decoding speedup at 128K context with near-baseline quality.
- LatentIndex: Cross-Layer Sharing with Layer-Specific Selection for Sparse Attention
LatentIndex shares continuous key representations across layers in sparse-attention indexers, improving recall by up to 3.28 points over IndexCache while cutting indexer-cache storage by 61.1%.