Summary (Overview)
- Problem: MoE models serving multi-turn agentic workloads face a "prefix-cache cliff" – reusable KV-cache prefixes grow across turns, but static HBM partitioning between expert weights and KV cache prevents reallocating memory, causing TTFT to spike sharply when cache capacity is exhausted.
- Key insight: Expert weights occupy most HBM (72 GiB per GPU for Qwen3-Next-80B on H100) despite only ~3B of 80B parameters being active per token; this memory can be dynamically repurposed for KV cache when needed.
- Solution: VAMP (Value-Aware Memory Partitioning) uses a three-way cost model comparing expert offloading, KV-cache eviction, and request preemption penalties, then uses CUDA Virtual Memory Management (VMM) to remap physical HBM pages between expert weights and KV cache at runtime.
- Results: On a 2,103-turn recorded SWE-bench replay, VAMP-15 reduces TTFT p90 from 26.1s to 1.10s (23.6×), increases request throughput by 20.7%, and raises prefix-cache hit rate from 0.19% to 89.8%, at a cost of +31.1% TPOT.
- Impact: VAMP expands the per-GPU KV-cache pool by up to 2.7× on H100, moving the capacity boundary (onset of the prefix-cache cliff) to significantly higher concurrency across both TP2 and TP4 configurations.
Introduction and Theoretical Foundation
Background
Large language model serving systems face a fundamental tension: model weights must remain resident in GPU HBM while KV-cache footprints grow linearly with batch size and context length. Traditional serving systems (e.g., vLLM) treat model placement and KV-cache management as separate decisions, relying on two implicit assumptions:
- Model weights occupy only a modest fraction of HBM, leaving ample room for KV cache
- KV-cache requirements remain within the preallocated pool bounds
Why These Assumptions Fail
Assumption 1 fails for MoE architectures: Mixture-of-Experts models structurally decouple model capacity from per-token compute cost. The total parameter count grows aggressively while active parameters stay constant:
| Model | Total Parameters | Active Parameters | Ratio |
|---|---|---|---|
| Mixtral-8x7B | 47B | ~13B | 3.6× |
| DeepSeek-V3 | 671B | 37B | 18× |
| Qwen3-Next-80B | 80B | 3B | ~27× |
Assumption 2 fails for multi-turn agentic workloads: Agentic trajectories average 157 turns and 32.7k tokens of context, with some runs exceeding 1M tokens. KV-cache utilization stays at 80–100%, causing repeated evictions and re-prefills.
The Prefix-Cache Cliff
When the reusable prefix-cache footprint exceeds KV-cache capacity:
- Cache eviction reduces prefix-cache hit rate
- Re-prefills inflate TTFT and consume serving capacity
- Sustained memory pressure leads to further preemptions
- This creates a threshold effect – TTFT remains stable until pool saturation, then rises sharply
This is not gradual degradation but an abrupt, threshold-driven collapse.
Methodology
VAMP Architecture
VAMP consists of three components:
- Value-Aware Action Selection – A three-way cost model comparing future penalties
- Dynamic Resource Orchestrator – Repartitions HBM between expert weights and KV cache
- Batching-Aware Expert Offloading – Overlaps weight transfers with computation
Cost Model
VAMP estimates the future cost of three actions when KV-cache allocation cannot be satisfied:
Cached-block eviction cost (): Uses a sliding-window prefix-cache hit rate and candidate age:
where is the number of tokens in reclaimed blocks, is the per-token re-prefill cost, and covers scheduling overhead.
Expert-offload cost (): Estimates per-step transfer cost over a decode-based accounting horizon :
where is the shard-equivalent size of the candidate expansion, is the number of local expert shards per layer per TP rank, and is a configured comparison coefficient.
Request preemption cost (): Based on the victim's last partially-filled block:
where is the number of additional uncached KV-cache blocks needed.
Combined action cost (): When expansion is not an exact multiple of the increment:
Dynamic HBM Repartitioning
- Uses CUDA VMM to unmap physical pages from expert groups and remap them into the KV-cache pool
- Expert tensors' virtual address ranges remain reserved; only physical backing is reassigned
- Expert groups are the minimum page-aligned remap unit (6 MiB per group for Qwen3-Next: two 3-MiB shards under TP2, four 1.5-MiB shards under TP4)
- Offloading bound: at most expert shards offloaded per layer, keeping at least resident
Batching-Aware Expert Offloading
- Each batched MoE-layer execution may require 262–473 of 512 routed experts (mean across 1,012 profiled steps), not just top-k=10
- VAMP loads the entire contiguous offloaded region into a staging buffer ahead of MoE execution
- Double buffering alternates between two staging buffers across consecutive layers, overlapping transfers with computation
Empirical Validation / Results
Experimental Setup
- Model: Qwen3-Next-80B-A3B-Instruct (512 routed experts/layer, top-k=10, 48 layers)
- Platforms: 2×H100 PCIe (TP2, memory-constrained) and 4×H200 NVLink (TP4)
- Baseline: vLLM 0.15.1 with prefix caching and chunked prefill
- Workloads: SWE-agent closed-loop, recorded trace replays (2,103 turns), LMCache traces, controlled in-house multi-turn
Measured KV-Cache Expansion (H100-PCIe)
| Config | Reclaimed (GiB/GPU) | Total KV pool (GiB/GPU) | KV pool/baseline |
|---|---|---|---|
| vLLM | — | 8.0 | 1.00× |
| VAMP-15 | 10.5 | 18.5 | 2.30× |
| VAMP-20 | 13.8 | 21.8 | 2.71× |
Recorded SWE-Agent Replay Results (n=5, mean±SD)
| Metric | vLLM | VAMP-15 | Difference |
|---|---|---|---|
| Request throughput (req/s) | 0.596 ± 0.007 | 0.720 ± 0.008 | +20.7% |
| TTFT p90 (s) | 26.1 ± 0.7 | 1.10 ± 0.17 | 23.6× lower |
| TTFT p99 (s) | 32.7 ± 2.7 | 3.48 ± 0.51 | 9.4× lower |
| Mean queue time (s) | 15.1 ± 0.5 | 0.021 ± 0.012 | -99.9% |
| Preemptions/1k turns | 562.7 ± 13.3 | 0.18 ± 0.41 | -99.9% |
| Prefix-cache hit rate (%) | 0.19 ± 0.03 | 89.8 ± 3.4 | +89.6 pp |
| TPOT (ms) | 170.9 ± 1.8 | 224.0 ± 2.0 | +31.1% |
Capacity Boundary Results
H100-PCIe: The baseline is already beyond capacity at c=16 (TTFT p99 = 22.1s), while VAMP-25 stays at 3.0–3.9s through c=28. The boundary moves with offloading ratio: VAMP-15 crosses at c=24, VAMP-20 at c=28, VAMP-25 doesn't cross within the range.
H200-NVL: All configurations stay within 9–11s through c=28. At c=32, baseline TTFT p99 rises to 92.8s while VAMP-20 remains at 9.1s (10.2× reduction).
Tradeoff Analysis
- VAMP-15 achieves the lowest TTFT p90 with throughput comparable to VAMP-10
- VAMP-20 provides comparable TTFT but higher TPOT and lower throughput
- The optimal offloading ratio depends on workload: at 128 LMCache sessions, VAMP-15 is better; at 64 sessions, VAMP-20 achieves lower TTFT p99 (1.30s vs 1.57s)
Theoretical and Practical Implications
Design Implications
-
Static memory partitioning is fundamentally inadequate for MoE models serving multi-turn workloads. The two-pool approach (weights vs. KV cache) creates artificial scarcity that triggers cascading evictions precisely when prefix reuse is highest.
-
Value-aware decisions require forward-looking cost estimates. The paper demonstrates that comparing the estimated future penalties of different actions (rather than just reacting to immediate pressure) is essential for optimal memory management.
-
CUDA VMM enables efficient cross-pool reallocation without data copying. Page remapping at runtime provides fine-grained control that was previously unavailable in serving frameworks.
-
Per-token sparsity ≠ small batched working set. Even with top-k=10, batched execution requires 262–473 distinct experts per layer, making LRU-style expert caching impractical.
Practical Implications
- VAMP demonstrates that MoE serving can operate far beyond static KV-cache capacity boundaries, extending the practical serving range for long-running agentic workloads
- The framework provides a tunable knob (maximum offloading ratio) that can be matched to workload characteristics
- The approach is complementary to other techniques: host-side KV-cache offloading, prefix-cache scheduling, and weight quantization
Conclusion
VAMP breaks the static boundary between expert-weight memory and the KV-cache pool in MoE serving, using CUDA VMM page remapping guided by a three-way cost model. Key results:
- 23.6× reduction in TTFT p90 (26.1s → 1.10s) on a 2,103-turn recorded trace
- 20.7% increase in request throughput with near-zero preemptions
- Up to 2.7× KV-cache pool expansion per GPU
- Prefix-cache hit rate improvement from 0.19% to 89.8%
The cost is increased TPOT (+31.1%) due to expert weight staging transfers.
Future directions identified by the authors:
- Restoring expert residency when KV-cache pressure subsides (currently one-way reallocation)
- Combining with host-side KV-cache offloading to extend coverage beyond the expanded capacity
- Matching offloading ratio to workload characteristics dynamically
- Exploring interconnect-aware scheduling to reduce contention between weight transfers and tensor-parallel traffic
Related papers
- ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI
ScientistTwo, a fully autonomous multi-agent framework, outperforms human state-of-the-art on 86 of 107 research problems, producing publication-quality papers with verified code without human intervention.
- DASC: Decay-Aware State Compression for Hybrid Linear-Attention Serving
DASC compresses recurrent state checkpoints by retaining only units with long retention horizons, achieving 2.63x compression, 42.6% lower TTFT, and 68.4% higher throughput with negligible quality loss.
- SWE-MeM: Learning Adaptive Memory Management for Long-Horizon Coding Agents
SWE-MeM trains agents to proactively compress their own context via a learned memory tool, achieving 60.2% on SWE-Bench Verified with a 30B model under a 32K budget, outperforming larger models and reducing token usage.