iS-KV: Online Low-Rank KV Cache Compression via Block-Incremental SVD
Summary (Overview)
-
Novel approach to KV cache compression: iS-KV retains every token position during long chain-of-thought (CoT) reasoning by storing older states in bounded-rank low-rank representations instead of evicting tokens, while keeping recent states at full precision.
-
Key insight on basis drift: The authors demonstrate that when the low-rank basis is updated for new tokens but old tokens' coordinates remain in the old basis, the stored history drifts substantially (drift of 1.55 for keys and 1.40 for values at step 256). iS-KV solves this by jointly updating both the basis and historical coordinates via block-incremental SVD.
-
Superior reasoning accuracy: On MATH-500, iS-KV achieves 82.6% accuracy at 4.06× compression on DeepSeek-R1-Distill-Llama-8B (close to the original's 83.6%) and 89.2% at 5.64× compression on Qwen3-8B, consistently outperforming token-eviction baselines (R-KV, SnapKV) at all six matched-memory operating points.
-
Token importance is unpredictable: The paper shows that ~30% of the 75% least-attended tokens at step 1024 receive strong attention again within the next 8K tokens on AIME benchmarks, growing to ~60% by 30K tokens—motivating the need to retain all positions.
-
Broad evaluation: The method is validated on reasoning benchmarks (MATH-500, AIME 2024/2025), long-input retrieval (RULER at 64K context), and efficiency measurements, with ablations confirming the critical role of online basis adaptation.
Introduction and Theoretical Foundation
Background and Motivation
Long chain-of-thought (CoT) reasoning significantly increases KV-cache memory during autoregressive decoding. Each generated token introduces new key and value states, causing the cache to grow linearly with decoding length. This creates a fundamental tension: reducing KV storage cost during decoding without losing information that later reasoning may need.
Two Key Observations
-
Old tokens can become important again: A token that receives little attention at the current step may receive strong attention thousands of tokens later. Once evicted, it cannot be recovered. The paper demonstrates this with R-KV, which reduces MATH-500 accuracy from 94.4% to 68.8% at ~5.6× compression on Qwen3-8B.
-
KV states have strong low-rank structure: Generated KV states exhibit significant low-rank structure (particularly keys before rotary positional embeddings), enabling compression without deletion.
The Basis Drift Problem
The central theoretical challenge: when the low-rank basis evolves as new tokens arrive, the coefficients of previously compressed tokens (computed in the old basis) must be updated accordingly. Otherwise, the same coefficients represent different vectors under the new basis, causing the stored history to drift from its true representation.
Theoretical Foundation
The optimization objective for compressing a set of rows with rank budget is:
The truncated SVD, denoted by , retains the leading singular components, where and contain orthonormal left/right singular vectors and is diagonal with retained singular values.
Methodology
Architecture Overview
iS-KV maintains two components for each K/V path:
- Exact recent window: The most recent tokens stored at original precision
- Low-rank history: Older tokens folded into compact representations with rank budget
Prefill Compression
-
Key protection: For each chunk and KV head , score by minimum cosine similarity: . Select up to lowest-scoring chunks per head as outlier chunks; the union of their positions plus a local suffix defines protected positions .
-
Value protection: includes all positions within tokens of any position in .
-
Initial factorization: Remaining rows are factorized via truncated SVD.
Block-Incremental Decoding Updates
For each regular update, given history with components and pending block , the goal is:
where . The update procedure:
- Compute residual basis: , then construct residual basis via QR and SVD
- Express history and block in augmented basis :
- Compute core SVD: , giving updated factors:
This reduces the SVD to a core, independent of history length .
Reconstruction and Storage
The reconstructed cache at position is:
Per-layer persistent storage cost:
where and are bytes per scalar for factors and exact rows, respectively.
Empirical Validation / Results
MATH-500 Reasoning Accuracy
Table 1: MATH-500 accuracy at matched persistent-KV footprints
| Method | Setting | KV ratio ↓ | Compression ↑ | Correct | Accuracy ↑ | Δ |
|---|---|---|---|---|---|---|
| DeepSeek-R1-Distill-Llama-8B | ||||||
| Original | - | 100.00% | 1.000× | 418/500 | 83.6 | 0.0 |
| iS-KV | r = 128 | 24.63% | 4.060× | 413/500 | 82.6 | -1.0 |
| R-KV | budget 748 | 24.63% | 4.060× | 391/500 | 78.2 | -5.4 |
| SnapKV | B = 636 | 24.64% | 4.059× | 351/500 | 70.1 | -13.5 |
| iS-KV | r = 96 | 20.79% | 4.811× | 409/500 | 81.8 | -1.8 |
| iS-KV | r = 64 | 16.90% | 5.917× | 391/500 | 78.2 | -5.4 |
| Qwen3-8B | ||||||
| Original | - | 100.00% | 1.000× | 472/500 | 94.4 | 0.0 |
| iS-KV | r = 128 | 21.36% | 4.681× | 460/500 | 92.0 | -2.4 |
| iS-KV | r = 96 | 17.72% | 5.644× | 446/500 | 89.2 | -5.2 |
| iS-KV | r = 64 | 14.05% | 7.118× | 439/500 | 87.8 | -6.6 |
| R-KV | budget 479 | 14.06% | 7.113× | 304/500 | 60.8 | -33.6 |
Key findings:
- iS-KV outperforms both eviction baselines at all six reported model–budget settings
- Accuracy degrades gradually with decreasing rank (4.4 points on DeepSeek, 4.2 on Qwen3 from rank 128→64)
- Gains over R-KV range from 5.4 to 5.6 percentage points at higher compression
AIME Results
On Qwen3-8B with 32,768-token limit:
- AIME 2024: At rank 192, iS-KV answers 23/30 correctly (76.7%), matching the original model, versus 43.3% for matched-memory R-KV
- AIME 2025: Rank 192 achieves 56.7%, above R-KV's 40.0% but below original's 63.3%
Comparison with OjaKV
The released OjaKV checkpoint achieves only 62.0% at ~1.30× compression, and all 118 generations at reduced ranks fell into repetitive loops reaching the length limit without EOS.
Ablation Study
Table 3: Component ablation on fixed 80-question MATH subset
| Variant | Low-rank K | Low-rank V | Online update | Compression | Correct | Accuracy (%) |
|---|---|---|---|---|---|---|
| Original | × | × | × | 1.00× | 67/80 | 83.8 |
| K-only | √ | × | √ | 1.64× | 68/80 | 85.0 |
| V-only | × | √ | √ | 1.64× | 65/80 | 81.2 |
| Frozen basis | √ | √ | × | 4.68× | 51/80 | 63.7 |
| iS-KV | √ | √ | √ | 4.61× | 65/80 | 81.2 |
Critical finding: Keeping bases fixed reduces accuracy by 17.5 percentage points, confirming the necessity of online basis adaptation.
Efficiency Results
- Generation time: iS-KV takes 47.45s vs 38.83s (original) at 8K input, with overhead decreasing from 22.2% to 8.8% as input length grows to 32K
- Memory savings: 33.7% at 8K, 44.2% at 16K, and 56.2% at 32K
Theoretical and Practical Implications
Theoretical Contributions
-
Basis-coordinate synchronization: The paper establishes that updating the basis without updating historical coordinates causes representation drift, and provides a mathematically grounded solution via block-incremental SVD that maintains consistency.
-
Computational efficiency: The core SVD is independent of history length , operating on a core, making online updates tractable for long sequences.
-
Storage scaling: Storage grows as with history length rather than , providing a tunable trade-off between fidelity and memory.
Practical Implications
-
Token eviction is fundamentally limited: The empirical demonstration that ~60% of low-attention tokens regain importance within 30K tokens challenges the core assumption of eviction-based methods.
-
Value compression is more challenging: Values decay more slowly in singular values and are harder to compress than keys, suggesting asymmetric rank allocations may be beneficial.
-
Long-prompt compression remains difficult: RULER results show more significant degradation on multi-value and multi-query retrieval at 64K context, indicating room for improvement in prompt-history compression.
Conclusion
iS-KV demonstrates that compressing a growing KV cache is fundamentally a problem of representing every token faithfully, not just choosing which tokens to keep. By jointly updating the low-rank basis and historical coordinates through block-incremental SVD, the method achieves near-original accuracy at 4–7× compression on long-horizon reasoning tasks, consistently outperforming token-eviction baselines at matched memory budgets.
Future directions suggested by the work include:
- Improving value compression, which remains more challenging than key compression
- Reducing the remaining accuracy gap at higher compression ratios
- Addressing long-prompt compression scenarios where the entire prompt exceeds the exact-recent window
- Potential integration with quantization and serving systems for additional memory savings
Related papers
- Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents
Verifier evolution lets agents self-improve without ground truth, but only anchor discipline—not detector lifecycle—prevents collapse into vacuous always-pass grading.
- Balancing Memory Pathways: Analyzing and Improving Memory Utilization in Hybrid LMs
Hybrid language models over-rely on attention and underuse recurrent memory, but an auxiliary loss forcing recurrent routing improves QA accuracy by up to 5.2% and agentic success rates.
- hacktrace: behavior-supervised detection of reward hacking during code generation
HACKTRACE detects reward hacking in coding agents from activations already computed during generation, cutting cheating from 85% to under 5% with only 8 ms overhead.