Summary (Overview)
- VFOLD is a novel, query-agnostic method for compressing the value cache of pre-trained Large Language Models (LLMs) by aligning and sharing value caches across adjacent layers, achieving 25% total KV cache reduction with negligible performance loss.
- The method exploits a weight-space symmetry in multi-headed attention to fold function-preserving linear maps (Canonical Correlation Analysis alignments and head permutations) directly into the and weight matrices, requiring no architectural changes, fine-tuning, or inference overhead until cache approximation occurs.
- Key finding: While both keys and values exhibit high Centered Kernel Alignment (CKA) similarity across layers, only values are amenable to cross-layer averaging. Averaging keys collapses performance (2.7% accuracy vs. 80.7% full cache), while averaging values retains 79.5% accuracy on GSM8K.
- VFOLD retains >98% of full-cache performance on RULER and LongBench benchmarks across three models (Llama-3.1-8B-Instruct, Mistral-Small-3.1-24B-Instruct, Qwen3-8B), outperforming prior cross-layer methods like MiniCache and CommonKV.
- VFOLD composes orthogonally with existing compression techniques (4-bit KIVI quantization, ThinK key pruning) to reach combined compression ratios (up to 4.27×) that neither method achieves alone.
Introduction and Theoretical Foundation
Background & Motivation: Long-context LLM decoding shifts the bottleneck from compute to memory, driven primarily by the KV cache. For a 16GB Llama-3.1-8B (fp16), a batch of 8 inputs at 16K tokens incurs a 16GB KV cache—matching the memory of the weights themselves. Existing compression methods fall into categories:
- Eviction: Selecting a subset of key/value pairs (e.g., H2O, StreamingLLM)
- Quantization: Storing keys/values at reduced bit width (e.g., KIVI)
- Low-rank representations: Storing rank-reduced caches (e.g., xKV, CommonKV)
- Cross-layer sharing: Sharing caches across layers (e.g., MiniCache)
Key Theoretical Insight: Prior work shows that while adjacent layers' key/value vectors have near-zero cosine similarity, they exhibit high CKA similarity (invariant to rotations and scaling), suggesting similarity exists only up to a change of basis. The paper exploits a linear invariance in attention: for a head with inputs , values , and attention matrix :
For any map , we can insert it between values and the output matrix:
Due to matrix associativity, can be folded into and into entirely offline. This invariance does not apply to keys due to RoPE, which would require maps to commute with rotation operations.
Methodology
1. Computing Alignment Maps (CCA): For a base layer and subsequent layer , Canonical Correlation Analysis (CCA) finds projections and that project each cache into a space of maximum correlation. The alignment map is defined as , applied to layer .
2. Optimal Head Pairing: Since dimensions from different KV heads cannot be mixed, heads are optimally paired between layers using the Hungarian algorithm. For every potential pairing , the alignment cost is:
The optimal permutation minimizes total cost:
3. Weight Folding: The permutation is applied to , , heads, and the block-diagonal map is applied to . The inverse operations (, ) are applied to . This yields a function-preserving reparameterization—the model computes exactly the same function.
4. Runtime Cache Merging: During decoding, the merged cache is computed as:
where . The current step always uses full value vectors; merging occurs only after attention computation. Protection is provided for 4 attention sinks and a sliding window of 128 recent tokens.
Empirical Validation / Results
Performance at 25% Total KV Reduction (50% value cache reduction):
RULER (16K) Results:
| Model | Full KV | MiniCache | MiniCache-V | CommonKV | VFOLD-NAIVE | VFOLD |
|---|---|---|---|---|---|---|
| Llama-3.1-8B-Inst. | 92.48 | 41.80 | 63.92 | 76.92 | 79.56 | 91.14 |
| Mistral-Small-24B | 95.94 | 28.34 | 83.45 | 94.95 | 91.13 | 95.12 |
| Qwen3-8B | 93.14 | 6.19 | 72.10 | 88.84 | 78.84 | 91.63 |
LongBench Results (average scores):
| Model | Full KV | MiniCache | MiniCache-V | CommonKV | VFOLD-NAIVE | VFOLD |
|---|---|---|---|---|---|---|
| Llama-3.1-8B-Inst. | 50.0 | 38.8 | 41.0 | 47.3 | 48.0 | 49.4 |
| Mistral-Small-24B | 55.3 | 44.6 | 49.8 | 55.2 | 53.3 | 54.7 |
| Qwen3-8B | 49.5 | 18.0 | 40.8 | 47.6 | 46.7 | 48.9 |
Efficiency Benchmarks (Llama-3.1-8B-Instruct, 8K context, A100-80G):
| Method | KV red. | KV mem. (GB) ↓ | TTFT (ms) ↓ | TPOT (ms) ↓ | Max batch ↑ |
|---|---|---|---|---|---|
| Full KV | 0% | 1.11 | 781 | 21.0 | 33 |
| MiniCache-V | 25% | 0.84 | 794 | 34.3 | 38 |
| CommonKV | 25% | 0.86 | 885 | 84.7 | 34 |
| VFOLD | 25% | 0.84 | 786 | 27.8 | 38 |
Composition with Other Methods (ThinK + VFOLD at 50% total reduction):
| Method | KV red. | Llama LB | Llama RULER | Mistral LB | Mistral RULER | Qwen LB | Qwen RULER |
|---|---|---|---|---|---|---|---|
| Full cache | 0% | 50.04 | 92.48 | 55.33 | 95.94 | 49.52 | 93.14 |
| ThinK | 25% | 49.65 | 88.90 | 52.53 | 92.44 | 49.28 | 93.05 |
| VFOLD | 25% | 49.40 | 91.14 | 54.72 | 95.12 | 48.92 | 91.63 |
| ThinK+VFOLD | 50% | 48.83 | 86.32 | 52.85 | 87.32 | 49.00 | 91.19 |
Key Results:
- VFOLD retains ≥98.3% of full-cache performance on every model and benchmark
- Alignment maps recover substantial performance: e.g., CWE task 2.2→75.3 for Llama-3.1-8B
- VFOLD adds minimal prefill overhead (786 vs. 781 ms TTFT) and lowest decoding overhead among compression methods
- Composition with 4-bit KIVI achieves 4.27× compression with performance within 0.3 points of VFOLD alone
Theoretical and Practical Implications
Theoretical Contributions:
- Key-Value Asymmetry: The paper provides empirical evidence and theoretical explanation for why values are more amenable to cross-layer merging than keys—key perturbation changes attention patterns amplified by softmax, while value perturbation is subject to linear combination across prior values.
- Symmetry Exploitation: Demonstrates that weight-space symmetries in attention can be exploited for cache compression without architectural modification, contrasting with low-rank methods that require "always-on" compression.
- Alignment Value: The CCA-based alignment maps substantially improve mergeability (cosine similarity from ~0.7 to ~0.85), showing that cross-layer similarity exists primarily in latent space rather than directly.
Practical Implications:
- VFOLD offers a drop-in solution for pre-trained models with no architectural changes, fine-tuning, or inference overhead
- The method enables larger batch sizes (+15% on A100-80G) at the same memory budget
- Composition with orthogonal compression methods (quantization, key pruning) allows reaching compression ratios unattainable by any single method
- Applicable to models with Grouped-Query Attention and Query-Key Normalization
Conclusion
VFOLD introduces a simple yet highly effective approach to value cache compression by exploiting attention weight symmetries to fold alignment maps into model weights offline. The method achieves 25% total KV cache reduction with >98% performance retention, outperforming prior cross-layer methods in both accuracy and efficiency. Key takeaways:
- Values, not keys, are the ideal target for cross-layer merging due to their role in attention and the availability of a linear invariance symmetry.
- Function-preserving reparameterization allows full-value computation during attention with compression applied only at cache storage, avoiding the overhead of always-on low-rank methods.
- Orthogonal composition with quantization (KIVI) and key pruning (ThinK) enables combined compression ratios up to 4.27×, addressing the growing memory bottleneck in long-context LLM decoding.
Future Directions: Extending the method to larger layer groups (explored in Appendix D), applying similar symmetry-aware folding to other cache dimensions, and investigating additional alignment map families beyond CCA.
Related papers
- TestJack: Should you trust the results in coding benchmarks? Agentic Coding Benchmarks Auditing via Evaluator Evolution
TESTJACK refutes 34.4% of benchmark-passing LLM coding trials by evolving tests from prompt requirements, cutting reported resolution rates from 50.6% to 33.2%.
- Complex Agents, Shallow Tests: Demystifying and Enhancing Test Adequacy of Agent Harness in the Wild
Agent harnesses remain substantially undertested, with less than half of LLM-dependent code covered, and HarnessTester's contract-aware test generation boosts coverage and mutation scores by up to 95% while finding 88 previously-unknown bugs.
- Gains and Collapse in On-Policy Distillation: A Reinforcement Learning Perspective
On-policy distillation improves sampling efficiency without expanding capability, and its collapse stems from reward hacking when teacher preferences misalign with response quality.