Full text not available for this paper
KV-Kaizen: Learning Context-Adaptive Cache Compression Choices
Summary (Overview)
- KV-Kaizen is a novel method for KV cache compression in LLMs that learns per-layer, context-adaptive choices along three orthogonal compression axes: depth (cross-layer cache sharing), rank (low-rank latent truncation via MLA), and precision (per-layer bit-width quantization).
- The method trains a lightweight selector (2.6–3.7M parameters, <0.05% of a 7B model) that reads the prompt once at pre-fill time and outputs a per-layer cache configuration satisfying a target compression budget, expressed in bits per token.
- Key finding: composing mild compression along multiple axes preserves accuracy far better than pushing any single axis aggressively. At 4× compression, KV-Kaizen matches uncompressed accuracy on models ≥7B parameters.
- KV-Kaizen composes multiplicatively with token eviction methods: combining 4× representation compression with SnapKV (keep 1/8) achieves 32× smaller decode-time cache on Qwen2.5-14B with negligible accuracy loss.
- The method supports budget-conditioned selectors (one model serving multiple compression targets) and demonstrates that compressed 7B models outperform smaller uncompressed models at matched cache sizes.
Introduction and Theoretical Foundation
Background: The KV Cache Bottleneck
Attention-based LLMs trade memory for compute: the KV cache stores key/value vectors for every prefilled or generated token, avoiding recomputation at each decoding step. However, this cache grows with sequence length and batch size, creating a memory bottleneck:
- Qwen2.5-14B uses 192 KiB per token of cache memory; serving 32 sequences of 8k tokens uses 48 GiB—more than the model weights in bf16.
- Decoding is memory-bandwidth-bound: every decoded token reads the entire KV cache, so cache size directly limits throughput.
Three Non-Eviction Compression Axes
The paper focuses on compression that does not drop tokens, distinguishing it from eviction-based approaches:
-
Depth (cross-layer attention): Layers share KV caches—an "anchor" layer produces a cache; subsequent layers "inherit" from the nearest anchor. The number of anchor layers determines effective depth.
-
Rank (MLA truncation): Multi-head latent attention (MLA) stores a shared latent representation plus a decoupled RoPE key . Truncating the latent to its first coordinates controls cache size. Non-MLA models are converted via CARE.
-
Precision (quantization): Cached tensors are stored at bits via symmetric uniform quantization (one scale per token), with straight-through gradient estimation during fine-tuning.
Key Insight
"Applying these structural changes aggressively quickly results in performance degradation."
The central hypothesis: individually mild compression along each axis, composed per-layer and adapted to context, achieves compression ratios no single axis reaches alone without accuracy loss.
Methodology
Action Space and Cost Model
Each layer chooses one action from the space:
where (bit-widths) and (kept latent widths). This yields 17 actions per layer.
The per-token cost of action is:
The compression ratio for a plan is:
where is the uncompressed bf16 non-MLA cache cost.
Selector Architecture and Training
- Selector: Two bidirectional transformer blocks ingest the backbone's token embeddings; representations are mean-pooled over time. One linear head per layer maps this shared feature to logits over .
- Selection: Gumbel-softmax with straight-through estimator (hard argmax forward, softmax backward), temperature annealed from 2.0 to 0.1.
- Initialization: Heads initialized so every layer's argmax is the least-compressed action (start from uncompressed model).
- Inference: Plain argmax; selector runs once before pre-fill (~3 ms, 0.20–3.3% of pre-fill time); configuration fixed for all generated tokens.
Training Objective
Standard language modeling loss plus a rate penalty:
where is layer 's straight-through one-hot selection vector. The hyperparameter is adaptively adjusted during training (raised when compression target is missed, relaxed when met), enabling one setting to reach all targets.
Budget-Conditioned Selectors
To serve multiple compression targets with one selector, the target is made an input via a zero-initialized FiLM layer on the pooled feature; training draws a random target per sequence.
Empirical Validation / Results
Main Results (Qwen2.5-7B)
Figure 3 shows mean IFEval + GSM8K accuracy vs. cache compression:
- 4× compression with KV-Kaizen matches uncompressed accuracy (0.679).
- Static (uniform per-layer) and random configurations degrade accuracy at matched budgets.
- KIVI (post-hoc 2-bit quantization) stays on the frontier up to ~5×; beyond that, only learned configurations work.
- Applying learned configurations to a frozen backbone collapses accuracy (must fine-tune under the configuration).
- Compressed 7B beats smaller uncompressed models: past 2×, a compressed 7B cache outperforms full-cache 3B, 1.5B, and 0.5B models.
Scaling with Model Size
Table 1 (mean IFEval + GSM8K accuracy differences at 4× compression):
| Approaches | Ref. | 3B | 7B | 14B | Mistral-7B | OLMo-3-7B |
|---|---|---|---|---|---|---|
| Uncompressed control | – | 0.673 | 0.679 | 0.723 | 0.349 | 0.508 |
| Precision (non-MLA) | −0.041 | −0.009 | +0.040 | +0.079 | +0.011 | |
| Depth (non-MLA) | −0.354 | −0.283 | −0.147 | −0.092 | −0.126 | |
| Precision + depth | −0.063 | −0.000 | +0.031 | +0.068 | +0.015 | |
| All three (MLA) | −0.169 | −0.028 | +0.021 | +0.027 | +0.000 |
Bold marks differences ≥ −0.053 (largest seed range).
Key patterns:
- Larger models tolerate compression better: gaps close from 3B to 14B.
- Depth alone is the weakest axis (must sacrifice 3/4 of layers to reach 4×).
- Combining axes doesn't add costs: precision+depth ≈ precision alone, while depth alone is far worse.
- Depth costs more on reasoning (GSM8K) than instruction following (IFEval).
Long-Context Evaluation and Eviction Composition (Table 2)
RULER at 16k on Qwen2.5-14B (QA accuracy):
| Configuration | ρ | NIAH multi-key | QA (SQuAD) |
|---|---|---|---|
| Uncompressed control | 1.0 | 0.956 | 0.543 |
| KVK, 4× (precision + depth) | 4.0 | 0.967 | 0.614 (+0.071) |
| + SnapKV, keep 0.25 | 16.0 | 0.967 | 0.560 (+0.017) |
| + SnapKV, keep 0.125 | 32.0 | 0.967 | 0.531 (−0.012) |
| KIVI 4-bit G=128 | 3.7 | 0.971 | 0.515 (−0.028) |
| SnapKV, keep 0.25 | 4.0 | 0.935 | 0.501 (−0.042) |
| PyramidKV, keep 0.25 | 4.0 | 0.949 | 0.485 (−0.057) |
Key result: KV-Kaizen at 4× beats eviction-only methods at matched cache size (0.614 vs 0.501 for SnapKV). Combining KV-Kaizen with eviction multiplies compression: 32× total with only −0.012 accuracy difference. KIVI + SnapKV is worse on both axes (29.5×, 0.464).
Budget-Conditioned Selectors (Table 3)
- Trained on {1×, 2×, 4×} with precision+depth: one selector covers all three, ahead of uncompressed control at 1×, only 0.067 behind single-target at 4×.
- Wider range [2×, 8×] with all three axes requires more training: 540k steps gets within 0.038/0.023/0.039 of single-target selectors at 2×/4×/8×.
- Too-wide target ranges degrade performance (selector stops following requests).
Hardware Benchmarking (Apple M3 Ultra)
Qwen2.5-14B with 4-bit weights, batch size 1:
- Composed compression (4-bit cache + MLA rank 128 + 12 groups of 4 layers): 152× total compression.
- Peak decode memory: 12.3 GB (bf16) → 0.17 GB (composed).
- Depth always helps decode throughput; MLA pays off at long sequences; 4-bit cache adds de-quantization compute cost but wins at long horizons.
Theoretical and Practical Implications
Why Composing Mild Compression Works
The paper's central theoretical contribution is demonstrating that compression axes are complementary rather than additive in their costs. Precision and rank degrade a layer "by degrees" (coarser values, narrower latent), while depth is a discrete change (layer computes from foreign keys/values). The selector naturally uses depth least because it's the most disruptive, but depth enables compression factors unreachable by precision alone (which floors at 8× due to 2-bit minimum).
Training Under the Configuration is Mandatory
"Applying a configuration to a model that never trained under it degrades it systematically."
This is a critical practical insight: post-hoc compression (even with good quantizers like KIVI) hits a ceiling around 5×; beyond that, the model must be fine-tuned under the target configuration.
Larger Models are More Compression-Tolerant
The consistent finding that larger models tolerate compression better supports a clear practical recommendation: pre-train large models, compress them afterward. Compressed 7B models outperform smaller uncompressed models at matched cache sizes, favoring the "train big, serve compressed" paradigm.
Selector Diversity and Deployment Considerations
- Some selectors return a single global configuration for all prompts (could be replaced by a static config); others adapt to context. This must be verified per selector.
- Budget-conditioned selectors enable deployment flexibility (one model, multiple cache budgets) at some accuracy cost.
Conclusion
Main Takeaways
- KV-Kaizen learns per-layer, context-adaptive cache configurations along depth, rank, and precision axes under a unified bit-per-token budget.
- Composing mild compression along multiple axes preserves accuracy at compression ratios (4×+) that no single axis achieves alone.
- 4× compression is essentially free for models ≥7B; compressed larger models beat smaller uncompressed ones.
- Compression composes multiplicatively with eviction (32× total on 14B at 16k context).
- Training under the serving configuration is essential; post-hoc methods plateau at ~5×.
Limitations and Future Work
- Specialized kernels needed to fully realize speed/memory benefits of mixed per-layer configurations.
- MLA conversion cost varies across model families; rank axis unavailable without it.
- Budget conditioning degrades with wider target ranges; finer control could enable per-deployment adaptation.
- Depth compression disproportionately hurts reasoning tasks; mechanisms to mitigate this are unexplored.
- The 32× result combines with eviction but was only demonstrated at 16k context; longer contexts remain untested.
Future Directions
- Finer-grained budget control for hardware-adaptive deployment.
- Specialized inference kernels for mixed-width/rank/depth configurations.
- Mitigating depth compression's impact on reasoning.
- Extending budget-conditioned selectors to wider target ranges without accuracy loss.
Related papers
- Identical Runs, Different Results: Benchmarking AI Coding Agents on Open-Weight Models
Identical runs of the same AI coding agent vary more than differences between agent-model pairings, so best-of-three with compliance checking beats single-run benchmarking.
- Cost-free Spectral Estimation for Adaptive Newton--Schulz in Matrix Optimizers
Newton–Schulz iterations expose spectral moments via free scalar reductions, enabling spectrum-adaptive routine selection that cuts polar error up to 90x and improves LLM pretraining loss.
- SOLAR: A State-Driven Online Learning Rate Scheduler for LLM Pretraining
SOLAR uses reinforcement learning to learn bounded residual corrections to a base learning-rate schedule, improving LLM pretraining perplexity across dense and MoE models up to 3B parameters.