Summary (Overview)

  • VFOLD is a novel, query-agnostic method for compressing the value cache of pre-trained Large Language Models (LLMs) by aligning and sharing value caches across adjacent layers, achieving 25% total KV cache reduction with negligible performance loss.
  • The method exploits a weight-space symmetry in multi-headed attention to fold function-preserving linear maps (Canonical Correlation Analysis alignments and head permutations) directly into the WVW_V and WOW_O weight matrices, requiring no architectural changes, fine-tuning, or inference overhead until cache approximation occurs.
  • Key finding: While both keys and values exhibit high Centered Kernel Alignment (CKA) similarity across layers, only values are amenable to cross-layer averaging. Averaging keys collapses performance (2.7% accuracy vs. 80.7% full cache), while averaging values retains 79.5% accuracy on GSM8K.
  • VFOLD retains >98% of full-cache performance on RULER and LongBench benchmarks across three models (Llama-3.1-8B-Instruct, Mistral-Small-3.1-24B-Instruct, Qwen3-8B), outperforming prior cross-layer methods like MiniCache and CommonKV.
  • VFOLD composes orthogonally with existing compression techniques (4-bit KIVI quantization, ThinK key pruning) to reach combined compression ratios (up to 4.27×) that neither method achieves alone.

Introduction and Theoretical Foundation

Background & Motivation: Long-context LLM decoding shifts the bottleneck from compute to memory, driven primarily by the KV cache. For a 16GB Llama-3.1-8B (fp16), a batch of 8 inputs at 16K tokens incurs a 16GB KV cache—matching the memory of the weights themselves. Existing compression methods fall into categories:

  • Eviction: Selecting a subset of key/value pairs (e.g., H2O, StreamingLLM)
  • Quantization: Storing keys/values at reduced bit width (e.g., KIVI)
  • Low-rank representations: Storing rank-reduced caches (e.g., xKV, CommonKV)
  • Cross-layer sharing: Sharing caches across layers (e.g., MiniCache)

Key Theoretical Insight: Prior work shows that while adjacent layers' key/value vectors have near-zero cosine similarity, they exhibit high CKA similarity (invariant to rotations and scaling), suggesting similarity exists only up to a change of basis. The paper exploits a linear invariance in attention: for a head hh with inputs XX, values Vh=XWVhV^h = XW_V^h, and attention matrix AhA^h:

Output of Head h=(AhVh)WOh∈Rn×dmodel(1)\text{Output of Head } h = (A^h V^h) W_O^h \in \mathbb{R}^{n \times d_{\text{model}}}\tag{1}

For any map Th∈GLdh(R)T_h \in \mathrm{GL}_{d_h}(\mathbb{R}), we can insert it between values and the output matrix:

[Ah(XWVh)Th][Th−1WOh](2)[A^h (X W_V^h) T_h] [T_h^{-1} W_O^h]\tag{2}

Due to matrix associativity, ThT_h can be folded into WVhW_V^h and Th−1T_h^{-1} into WOhW_O^h entirely offline. This invariance does not apply to keys due to RoPE, which would require maps to commute with rotation operations.

Methodology

1. Computing Alignment Maps (CCA): For a base layer ℓ\ell and subsequent layer ℓ+1\ell+1, Canonical Correlation Analysis (CCA) finds projections UℓU_\ell and Uℓ+1U_{\ell+1} that project each cache into a space of maximum correlation. The alignment map is defined as T=Uℓ+1Uℓ−1T = U_{\ell+1} U_\ell^{-1}, applied to layer ℓ+1\ell+1.

2. Optimal Head Pairing: Since dimensions from different KV heads cannot be mixed, heads are optimally paired between layers using the Hungarian algorithm. For every potential pairing (i,j)(i,j), the alignment cost is:

Cij=cost(i,j)=∣∣VℓhiTij−1−Vℓ+1hj∣∣F(3)C_{ij} = \text{cost}(i,j) = ||V_\ell^{h_i} T_{ij}^{-1} - V_{\ell+1}^{h_j}||_F\tag{3}

The optimal permutation π\pi minimizes total cost:

π=arg⁡min⁡Πnkv∑i=1nkvCi,π(i)(4)\pi = \arg\min_{\Pi_{n_{\text{kv}}}} \sum_{i=1}^{n_{\text{kv}}} C_{i,\pi(i)}\tag{4}

3. Weight Folding: The permutation PP is applied to WQW_Q, WKW_K, WVW_V heads, and the block-diagonal map T=blockdiag(T1,…,Tnkv)T = \text{blockdiag}(T_1, \dots, T_{n_{\text{kv}}}) is applied to WVW_V. The inverse operations (T−1T^{-1}, PTP^T) are applied to WOW_O. This yields a function-preserving reparameterization—the model computes exactly the same function.

4. Runtime Cache Merging: During decoding, the merged cache is computed as:

Vmerge=12(Vℓ+F(Vℓ+1))V_{\text{merge}} = \frac{1}{2}(V_\ell + F(V_{\ell+1}))

where F(V)=VPTF(V) = VPT. The current step always uses full value vectors; merging occurs only after attention computation. Protection is provided for 4 attention sinks and a sliding window of 128 recent tokens.

Empirical Validation / Results

Performance at 25% Total KV Reduction (50% value cache reduction):

RULER (16K) Results:

ModelFull KVMiniCacheMiniCache-VCommonKVVFOLD-NAIVEVFOLD
Llama-3.1-8B-Inst.92.4841.8063.9276.9279.5691.14
Mistral-Small-24B95.9428.3483.4594.9591.1395.12
Qwen3-8B93.146.1972.1088.8478.8491.63

LongBench Results (average scores):

ModelFull KVMiniCacheMiniCache-VCommonKVVFOLD-NAIVEVFOLD
Llama-3.1-8B-Inst.50.038.841.047.348.049.4
Mistral-Small-24B55.344.649.855.253.354.7
Qwen3-8B49.518.040.847.646.748.9

Efficiency Benchmarks (Llama-3.1-8B-Instruct, 8K context, A100-80G):

MethodKV red.KV mem. (GB) ↓TTFT (ms) ↓TPOT (ms) ↓Max batch ↑
Full KV0%1.1178121.033
MiniCache-V25%0.8479434.338
CommonKV25%0.8688584.734
VFOLD25%0.8478627.838

Composition with Other Methods (ThinK + VFOLD at 50% total reduction):

MethodKV red.Llama LBLlama RULERMistral LBMistral RULERQwen LBQwen RULER
Full cache0%50.0492.4855.3395.9449.5293.14
ThinK25%49.6588.9052.5392.4449.2893.05
VFOLD25%49.4091.1454.7295.1248.9291.63
ThinK+VFOLD50%48.8386.3252.8587.3249.0091.19

Key Results:

  • VFOLD retains ≥98.3% of full-cache performance on every model and benchmark
  • Alignment maps recover substantial performance: e.g., CWE task 2.2→75.3 for Llama-3.1-8B
  • VFOLD adds minimal prefill overhead (786 vs. 781 ms TTFT) and lowest decoding overhead among compression methods
  • Composition with 4-bit KIVI achieves 4.27× compression with performance within 0.3 points of VFOLD alone

Theoretical and Practical Implications

Theoretical Contributions:

  1. Key-Value Asymmetry: The paper provides empirical evidence and theoretical explanation for why values are more amenable to cross-layer merging than keys—key perturbation changes attention patterns amplified by softmax, while value perturbation is subject to linear combination across prior values.
  2. Symmetry Exploitation: Demonstrates that weight-space symmetries in attention can be exploited for cache compression without architectural modification, contrasting with low-rank methods that require "always-on" compression.
  3. Alignment Value: The CCA-based alignment maps substantially improve mergeability (cosine similarity from ~0.7 to ~0.85), showing that cross-layer similarity exists primarily in latent space rather than directly.

Practical Implications:

  • VFOLD offers a drop-in solution for pre-trained models with no architectural changes, fine-tuning, or inference overhead
  • The method enables larger batch sizes (+15% on A100-80G) at the same memory budget
  • Composition with orthogonal compression methods (quantization, key pruning) allows reaching compression ratios unattainable by any single method
  • Applicable to models with Grouped-Query Attention and Query-Key Normalization

Conclusion

VFOLD introduces a simple yet highly effective approach to value cache compression by exploiting attention weight symmetries to fold alignment maps into model weights offline. The method achieves 25% total KV cache reduction with >98% performance retention, outperforming prior cross-layer methods in both accuracy and efficiency. Key takeaways:

  1. Values, not keys, are the ideal target for cross-layer merging due to their role in attention and the availability of a linear invariance symmetry.
  2. Function-preserving reparameterization allows full-value computation during attention with compression applied only at cache storage, avoiding the overhead of always-on low-rank methods.
  3. Orthogonal composition with quantization (KIVI) and key pruning (ThinK) enables combined compression ratios up to 4.27×, addressing the growing memory bottleneck in long-context LLM decoding.

Future Directions: Extending the method to larger layer groups (explored in Appendix D), applying similar symmetry-aware folding to other cache dimensions, and investigating additional alignment map families beyond CCA.

Related papers