# VFold: Symmetry-Aware Cross-Layer Value Cache Compression

> VFOLD compresses LLM value caches by folding cross-layer alignment maps into attention weights, achieving 25% KV reduction with over 98% performance retention.

- **Source:** [arXiv](https://arxiv.org/abs/2610.12338)
- **Published:** 2026-10-10
- **Permalink:** https://picx.dev/p/zJFooq
- **Whiteboard:** https://picx.dev/p/zJFooq/image

## Summary

## Summary (Overview)

- **VFOLD** is a novel, query-agnostic method for compressing the value cache of pre-trained Large Language Models (LLMs) by aligning and sharing value caches across adjacent layers, achieving 25% total KV cache reduction with negligible performance loss.
- The method exploits a **weight-space symmetry in multi-headed attention** to fold function-preserving linear maps (Canonical Correlation Analysis alignments and head permutations) directly into the $W_V$ and $W_O$ weight matrices, requiring no architectural changes, fine-tuning, or inference overhead until cache approximation occurs.
- **Key finding**: While both keys and values exhibit high Centered Kernel Alignment (CKA) similarity across layers, only values are amenable to cross-layer averaging. Averaging keys collapses performance (2.7% accuracy vs. 80.7% full cache), while averaging values retains 79.5% accuracy on GSM8K.
- VFOLD retains **>98% of full-cache performance** on RULER and LongBench benchmarks across three models (Llama-3.1-8B-Instruct, Mistral-Small-3.1-24B-Instruct, Qwen3-8B), outperforming prior cross-layer methods like MiniCache and CommonKV.
- VFOLD **composes orthogonally** with existing compression techniques (4-bit KIVI quantization, ThinK key pruning) to reach combined compression ratios (up to 4.27×) that neither method achieves alone.

## Introduction and Theoretical Foundation

**Background & Motivation**: Long-context LLM decoding shifts the bottleneck from compute to memory, driven primarily by the KV cache. For a 16GB Llama-3.1-8B (fp16), a batch of 8 inputs at 16K tokens incurs a 16GB KV cache—matching the memory of the weights themselves. Existing compression methods fall into categories:

- **Eviction**: Selecting a subset of key/value pairs (e.g., H2O, StreamingLLM)
- **Quantization**: Storing keys/values at reduced bit width (e.g., KIVI)
- **Low-rank representations**: Storing rank-reduced caches (e.g., xKV, CommonKV)
- **Cross-layer sharing**: Sharing caches across layers (e.g., MiniCache)

**Key Theoretical Insight**: Prior work shows that while adjacent layers' key/value vectors have near-zero cosine similarity, they exhibit high CKA similarity (invariant to rotations and scaling), suggesting similarity exists only up to a change of basis. The paper exploits a **linear invariance in attention**: for a head $h$ with inputs $X$, values $V^h = XW_V^h$, and attention matrix $A^h$:

$$
\text{Output of Head } h = (A^h V^h) W_O^h \in \mathbb{R}^{n \times d_{\text{model}}}\tag{1}
$$

For any map $T_h \in \mathrm{GL}_{d_h}(\mathbb{R})$, we can insert it between values and the output matrix:

$$
[A^h (X W_V^h) T_h] [T_h^{-1} W_O^h]\tag{2}
$$

Due to matrix associativity, $T_h$ can be folded into $W_V^h$ and $T_h^{-1}$ into $W_O^h$ entirely offline. This invariance **does not apply to keys** due to RoPE, which would require maps to commute with rotation operations.

## Methodology

**1. Computing Alignment Maps (CCA)**:
For a base layer $\ell$ and subsequent layer $\ell+1$, Canonical Correlation Analysis (CCA) finds projections $U_\ell$ and $U_{\ell+1}$ that project each cache into a space of maximum correlation. The alignment map is defined as $T = U_{\ell+1} U_\ell^{-1}$, applied to layer $\ell+1$.

**2. Optimal Head Pairing**: Since dimensions from different KV heads cannot be mixed, heads are optimally paired between layers using the Hungarian algorithm. For every potential pairing $(i,j)$, the alignment cost is:

$$
C_{ij} = \text{cost}(i,j) = ||V_\ell^{h_i} T_{ij}^{-1} - V_{\ell+1}^{h_j}||_F\tag{3}
$$

The optimal permutation $\pi$ minimizes total cost:

$$
\pi = \arg\min_{\Pi_{n_{\text{kv}}}} \sum_{i=1}^{n_{\text{kv}}} C_{i,\pi(i)}\tag{4}
$$

**3. Weight Folding**: The permutation $P$ is applied to $W_Q$, $W_K$, $W_V$ heads, and the block-diagonal map $T = \text{blockdiag}(T_1, \dots, T_{n_{\text{kv}}})$ is applied to $W_V$. The inverse operations ($T^{-1}$, $P^T$) are applied to $W_O$. This yields a **function-preserving reparameterization**—the model computes exactly the same function.

**4. Runtime Cache Merging**: During decoding, the merged cache is computed as:

$$
V_{\text{merge}} = \frac{1}{2}(V_\ell + F(V_{\ell+1}))
$$

where $F(V) = VPT$. The current step always uses full value vectors; merging occurs only after attention computation. Protection is provided for 4 attention sinks and a sliding window of 128 recent tokens.

## Empirical Validation / Results

**Performance at 25% Total KV Reduction** (50% value cache reduction):

**RULER (16K) Results**:

| Model | Full KV | MiniCache | MiniCache-V | CommonKV | VFOLD-NAIVE | **VFOLD** |
|-------|---------|-----------|-------------|----------|-------------|-----------|
| Llama-3.1-8B-Inst. | 92.48 | 41.80 | 63.92 | 76.92 | 79.56 | **91.14** |
| Mistral-Small-24B | 95.94 | 28.34 | 83.45 | 94.95 | 91.13 | **95.12** |
| Qwen3-8B | 93.14 | 6.19 | 72.10 | 88.84 | 78.84 | **91.63** |

**LongBench Results** (average scores):

| Model | Full KV | MiniCache | MiniCache-V | CommonKV | VFOLD-NAIVE | **VFOLD** |
|-------|---------|-----------|-------------|----------|-------------|-----------|
| Llama-3.1-8B-Inst. | 50.0 | 38.8 | 41.0 | 47.3 | 48.0 | **49.4** |
| Mistral-Small-24B | 55.3 | 44.6 | 49.8 | 55.2 | 53.3 | **54.7** |
| Qwen3-8B | 49.5 | 18.0 | 40.8 | 47.6 | 46.7 | **48.9** |

**Efficiency Benchmarks** (Llama-3.1-8B-Instruct, 8K context, A100-80G):

| Method | KV red. | KV mem. (GB) ↓ | TTFT (ms) ↓ | TPOT (ms) ↓ | Max batch ↑ |
|--------|---------|----------------|-------------|-------------|-------------|
| Full KV | 0% | 1.11 | 781 | 21.0 | 33 |
| MiniCache-V | 25% | 0.84 | 794 | 34.3 | 38 |
| CommonKV | 25% | 0.86 | 885 | 84.7 | 34 |
| **VFOLD** | 25% | **0.84** | **786** | **27.8** | **38** |

**Composition with Other Methods** (ThinK + VFOLD at 50% total reduction):

| Method | KV red. | Llama LB | Llama RULER | Mistral LB | Mistral RULER | Qwen LB | Qwen RULER |
|--------|---------|----------|-------------|------------|---------------|---------|------------|
| Full cache | 0% | 50.04 | 92.48 | 55.33 | 95.94 | 49.52 | 93.14 |
| ThinK | 25% | 49.65 | 88.90 | 52.53 | 92.44 | 49.28 | 93.05 |
| VFOLD | 25% | 49.40 | 91.14 | 54.72 | 95.12 | 48.92 | 91.63 |
| **ThinK+VFOLD** | **50%** | **48.83** | **86.32** | **52.85** | **87.32** | **49.00** | **91.19** |

**Key Results**:
- VFOLD retains ≥98.3% of full-cache performance on every model and benchmark
- Alignment maps recover substantial performance: e.g., CWE task 2.2→75.3 for Llama-3.1-8B
- VFOLD adds minimal prefill overhead (786 vs. 781 ms TTFT) and lowest decoding overhead among compression methods
- Composition with 4-bit KIVI achieves 4.27× compression with performance within 0.3 points of VFOLD alone

## Theoretical and Practical Implications

**Theoretical Contributions**:
1. **Key-Value Asymmetry**: The paper provides empirical evidence and theoretical explanation for why values are more amenable to cross-layer merging than keys—key perturbation changes attention patterns amplified by softmax, while value perturbation is subject to linear combination across prior values.
2. **Symmetry Exploitation**: Demonstrates that weight-space symmetries in attention can be exploited for cache compression without architectural modification, contrasting with low-rank methods that require "always-on" compression.
3. **Alignment Value**: The CCA-based alignment maps substantially improve mergeability (cosine similarity from ~0.7 to ~0.85), showing that cross-layer similarity exists primarily in latent space rather than directly.

**Practical Implications**:
- VFOLD offers a **drop-in solution** for pre-trained models with no architectural changes, fine-tuning, or inference overhead
- The method enables **larger batch sizes** (+15% on A100-80G) at the same memory budget
- Composition with orthogonal compression methods (quantization, key pruning) allows reaching compression ratios unattainable by any single method
- Applicable to models with Grouped-Query Attention and Query-Key Normalization

## Conclusion

VFOLD introduces a simple yet highly effective approach to value cache compression by exploiting attention weight symmetries to fold alignment maps into model weights offline. The method achieves 25% total KV cache reduction with >98% performance retention, outperforming prior cross-layer methods in both accuracy and efficiency. Key takeaways:

1. **Values, not keys, are the ideal target for cross-layer merging** due to their role in attention and the availability of a linear invariance symmetry.
2. **Function-preserving reparameterization** allows full-value computation during attention with compression applied only at cache storage, avoiding the overhead of always-on low-rank methods.
3. **Orthogonal composition** with quantization (KIVI) and key pruning (ThinK) enables combined compression ratios up to 4.27×, addressing the growing memory bottleneck in long-context LLM decoding.

**Future Directions**: Extending the method to larger layer groups (explored in Appendix D), applying similar symmetry-aware folding to other cache dimensions, and investigating additional alignment map families beyond CCA.

---

_Markdown view of https://picx.dev/p/zJFooq, served by PicX — AI-generated visual whiteboard summaries of research papers._
