# KV-Kaizen: Learning Context-Adaptive Cache Compression Choices

> KV-Kaizen learns per-layer cache compression choices across depth, rank, and precision, achieving 4x compression with no accuracy loss on models 7B and larger.

- **Source:** [arXiv](https://arxiv.org/abs/2609.37988)
- **Published:** 2026-10-03
- **Permalink:** https://picx.dev/p/FDpc0d
- **Whiteboard:** https://picx.dev/p/FDpc0d/image

## Summary

# KV-Kaizen: Learning Context-Adaptive Cache Compression Choices

## Summary (Overview)

- **KV-Kaizen** is a novel method for KV cache compression in LLMs that learns per-layer, context-adaptive choices along three orthogonal compression axes: **depth** (cross-layer cache sharing), **rank** (low-rank latent truncation via MLA), and **precision** (per-layer bit-width quantization).
- The method trains a lightweight selector (2.6–3.7M parameters, <0.05% of a 7B model) that reads the prompt once at pre-fill time and outputs a per-layer cache configuration satisfying a target compression budget, expressed in bits per token.
- Key finding: **composing mild compression along multiple axes** preserves accuracy far better than pushing any single axis aggressively. At 4× compression, KV-Kaizen matches uncompressed accuracy on models ≥7B parameters.
- KV-Kaizen **composes multiplicatively with token eviction methods**: combining 4× representation compression with SnapKV (keep 1/8) achieves **32× smaller decode-time cache** on Qwen2.5-14B with negligible accuracy loss.
- The method supports **budget-conditioned selectors** (one model serving multiple compression targets) and demonstrates that compressed 7B models outperform smaller uncompressed models at matched cache sizes.

## Introduction and Theoretical Foundation

### Background: The KV Cache Bottleneck

Attention-based LLMs trade memory for compute: the KV cache stores key/value vectors for every prefilled or generated token, avoiding recomputation at each decoding step. However, this cache grows with sequence length and batch size, creating a memory bottleneck:

- Qwen2.5-14B uses **192 KiB per token** of cache memory; serving 32 sequences of 8k tokens uses 48 GiB—more than the model weights in bf16.
- Decoding is **memory-bandwidth-bound**: every decoded token reads the entire KV cache, so cache size directly limits throughput.

### Three Non-Eviction Compression Axes

The paper focuses on compression that **does not drop tokens**, distinguishing it from eviction-based approaches:

1. **Depth (cross-layer attention)**: Layers share KV caches—an "anchor" layer produces a cache; subsequent layers "inherit" from the nearest anchor. The number of anchor layers determines effective depth.

2. **Rank (MLA truncation)**: Multi-head latent attention (MLA) stores a shared latent representation $c^{KV}_l \in \mathbb{R}^r$ plus a decoupled RoPE key $k^R_l \in \mathbb{R}^{d_R}$. Truncating the latent to its first $d \leq r$ coordinates controls cache size. Non-MLA models are converted via CARE.

3. **Precision (quantization)**: Cached tensors are stored at $b \leq 16$ bits via symmetric uniform quantization (one scale per token), with straight-through gradient estimation during fine-tuning.

### Key Insight

> "Applying these structural changes aggressively quickly results in performance degradation."

The central hypothesis: **individually mild** compression along each axis, **composed** per-layer and adapted to context, achieves compression ratios no single axis reaches alone without accuracy loss.

## Methodology

### Action Space and Cost Model

Each layer $l$ chooses one action $a_l$ from the space:

$$\mathcal{A} = \{\text{inherit}\} \cup \{(d, b) : d \in \mathcal{D}, b \in \mathcal{B}\}$$

where $\mathcal{B} = \{2, 4, 8, 16\}$ (bit-widths) and $\mathcal{D} = \{r/8, r/4, r/2, r\}$ (kept latent widths). This yields **17 actions per layer**.

The per-token cost of action $a$ is:

$$c(a) = \begin{cases} 0 & a = \text{inherit} \\ (d + d_R) \cdot b & a = (d, b) \end{cases} \quad \text{(bits per token)}$$

The compression ratio for a plan $a_{1:L}$ is:

$$\rho(a_{1:L}) = \frac{C_0}{\sum_{l=1}^L c(a_l)}, \quad C_0 = L \cdot 2 n_{kv} d_h \cdot 16$$

where $C_0$ is the uncompressed bf16 non-MLA cache cost.

### Selector Architecture and Training

- **Selector**: Two bidirectional transformer blocks ingest the backbone's token embeddings; representations are mean-pooled over time. One linear head per layer maps this shared feature to logits over $\mathcal{A}$.
- **Selection**: Gumbel-softmax with straight-through estimator (hard argmax forward, softmax backward), temperature annealed from 2.0 to 0.1.
- **Initialization**: Heads initialized so every layer's argmax is the least-compressed action (start from uncompressed model).
- **Inference**: Plain argmax; selector runs once before pre-fill (~3 ms, 0.20–3.3% of pre-fill time); configuration fixed for all generated tokens.

### Training Objective

Standard language modeling loss plus a rate penalty:

$$\mathcal{L}_{\text{rate}} = \beta \cdot \mathbb{E}_i \left| \frac{1}{\rho_i} - \frac{1}{\rho^\star} \right|, \quad \frac{1}{\rho_i} = \frac{1}{C_0} \sum_l \sum_{a \in \mathcal{A}} w^{(i)}_{l,a} c(a)$$

where $w^{(i)}_l$ is layer $l$'s straight-through one-hot selection vector. The hyperparameter $\beta$ is **adaptively adjusted** during training (raised when compression target is missed, relaxed when met), enabling one setting to reach all targets.

### Budget-Conditioned Selectors

To serve multiple compression targets with one selector, the target $\rho^\star$ is made an input via a zero-initialized FiLM layer on the pooled feature; training draws a random target per sequence.

## Empirical Validation / Results

### Main Results (Qwen2.5-7B)

**Figure 3** shows mean IFEval + GSM8K accuracy vs. cache compression:

- **4× compression with KV-Kaizen matches uncompressed accuracy** (0.679).
- Static (uniform per-layer) and random configurations degrade accuracy at matched budgets.
- KIVI (post-hoc 2-bit quantization) stays on the frontier up to ~5×; beyond that, only learned configurations work.
- Applying learned configurations to a *frozen* backbone collapses accuracy (must fine-tune under the configuration).
- Compressed 7B beats smaller uncompressed models: past 2×, a compressed 7B cache outperforms full-cache 3B, 1.5B, and 0.5B models.

### Scaling with Model Size

**Table 1** (mean IFEval + GSM8K accuracy differences at 4× compression):

| Approaches | Ref. | 3B | 7B | 14B | Mistral-7B | OLMo-3-7B |
|---|---|---|---|---|---|---|
| Uncompressed control | – | 0.673 | 0.679 | 0.723 | 0.349 | 0.508 |
| Precision (non-MLA) | | −0.041 | −0.009 | **+0.040** | **+0.079** | **+0.011** |
| Depth (non-MLA) | | −0.354 | −0.283 | −0.147 | −0.092 | −0.126 |
| Precision + depth | | −0.063 | **−0.000** | **+0.031** | **+0.068** | **+0.015** |
| All three (MLA) | | −0.169 | −0.028 | **+0.021** | **+0.027** | **+0.000** |

*Bold marks differences ≥ −0.053 (largest seed range).*

Key patterns:
- **Larger models tolerate compression better**: gaps close from 3B to 14B.
- **Depth alone is the weakest axis** (must sacrifice 3/4 of layers to reach 4×).
- **Combining axes doesn't add costs**: precision+depth ≈ precision alone, while depth alone is far worse.
- Depth costs more on reasoning (GSM8K) than instruction following (IFEval).

### Long-Context Evaluation and Eviction Composition (Table 2)

RULER at 16k on Qwen2.5-14B (QA accuracy):

| Configuration | ρ | NIAH multi-key | QA (SQuAD) |
|---|---|---|---|
| Uncompressed control | 1.0 | 0.956 | 0.543 |
| **KVK, 4× (precision + depth)** | **4.0** | **0.967** | **0.614 (+0.071)** |
| + SnapKV, keep 0.25 | 16.0 | 0.967 | 0.560 (+0.017) |
| + SnapKV, keep 0.125 | **32.0** | 0.967 | 0.531 (−0.012) |
| KIVI 4-bit G=128 | 3.7 | 0.971 | 0.515 (−0.028) |
| SnapKV, keep 0.25 | 4.0 | 0.935 | 0.501 (−0.042) |
| PyramidKV, keep 0.25 | 4.0 | 0.949 | 0.485 (−0.057) |

**Key result**: KV-Kaizen at 4× beats eviction-only methods at matched cache size (0.614 vs 0.501 for SnapKV). Combining KV-Kaizen with eviction multiplies compression: **32× total** with only −0.012 accuracy difference. KIVI + SnapKV is worse on both axes (29.5×, 0.464).

### Budget-Conditioned Selectors (Table 3)

- Trained on {1×, 2×, 4×} with precision+depth: one selector covers all three, ahead of uncompressed control at 1×, only 0.067 behind single-target at 4×.
- Wider range [2×, 8×] with all three axes requires more training: 540k steps gets within 0.038/0.023/0.039 of single-target selectors at 2×/4×/8×.
- Too-wide target ranges degrade performance (selector stops following requests).

### Hardware Benchmarking (Apple M3 Ultra)

Qwen2.5-14B with 4-bit weights, batch size 1:

- Composed compression (4-bit cache + MLA rank 128 + 12 groups of 4 layers): **152× total compression**.
- Peak decode memory: 12.3 GB (bf16) → 0.17 GB (composed).
- Depth always helps decode throughput; MLA pays off at long sequences; 4-bit cache adds de-quantization compute cost but wins at long horizons.

## Theoretical and Practical Implications

### Why Composing Mild Compression Works

The paper's central theoretical contribution is demonstrating that **compression axes are complementary rather than additive in their costs**. Precision and rank degrade a layer "by degrees" (coarser values, narrower latent), while depth is a discrete change (layer computes from foreign keys/values). The selector naturally uses depth least because it's the most disruptive, but depth enables compression factors unreachable by precision alone (which floors at 8× due to 2-bit minimum).

### Training Under the Configuration is Mandatory

> "Applying a configuration to a model that never trained under it degrades it systematically."

This is a critical practical insight: post-hoc compression (even with good quantizers like KIVI) hits a ceiling around 5×; beyond that, the model must be fine-tuned under the target configuration.

### Larger Models are More Compression-Tolerant

The consistent finding that larger models tolerate compression better supports a clear practical recommendation: **pre-train large models, compress them afterward**. Compressed 7B models outperform smaller uncompressed models at matched cache sizes, favoring the "train big, serve compressed" paradigm.

### Selector Diversity and Deployment Considerations

- Some selectors return a single global configuration for all prompts (could be replaced by a static config); others adapt to context. This must be verified per selector.
- Budget-conditioned selectors enable deployment flexibility (one model, multiple cache budgets) at some accuracy cost.

## Conclusion

### Main Takeaways

1. **KV-Kaizen** learns per-layer, context-adaptive cache configurations along depth, rank, and precision axes under a unified bit-per-token budget.
2. **Composing mild compression along multiple axes** preserves accuracy at compression ratios (4×+) that no single axis achieves alone.
3. **4× compression is essentially free** for models ≥7B; compressed larger models beat smaller uncompressed ones.
4. **Compression composes multiplicatively with eviction** (32× total on 14B at 16k context).
5. **Training under the serving configuration is essential**; post-hoc methods plateau at ~5×.

### Limitations and Future Work

- **Specialized kernels** needed to fully realize speed/memory benefits of mixed per-layer configurations.
- **MLA conversion cost** varies across model families; rank axis unavailable without it.
- **Budget conditioning degrades** with wider target ranges; finer control could enable per-deployment adaptation.
- Depth compression disproportionately hurts reasoning tasks; mechanisms to mitigate this are unexplored.
- The 32× result combines with eviction but was only demonstrated at 16k context; longer contexts remain untested.

### Future Directions

- Finer-grained budget control for hardware-adaptive deployment.
- Specialized inference kernels for mixed-width/rank/depth configurations.
- Mitigating depth compression's impact on reasoning.
- Extending budget-conditioned selectors to wider target ranges without accuracy loss.

---

_Markdown view of https://picx.dev/p/FDpc0d, served by PicX — AI-generated visual whiteboard summaries of research papers._
