# LatentIndex: Cross-Layer Sharing with Layer-Specific Selection for Sparse Attention

> LatentIndex shares continuous key representations across layers in sparse-attention indexers, improving recall by up to 3.28 points over IndexCache while cutting indexer-cache storage by 61.1%.

- **Source:** [arXiv](https://arxiv.org/abs/2610.04635)
- **Published:** 2026-10-10
- **Permalink:** https://picx.dev/p/VoMFGs
- **Whiteboard:** https://picx.dev/p/VoMFGs/image

## Summary

## Summary (Overview)

- **LatentIndex** extends the latent-sharing principle of Multi-head Latent Attention (MLA) to sparse-attention indexers, enabling cross-layer sharing of continuous key representations while preserving layer-specific token selection.
.
 The key innovation is replacing per-layer key caches with a shared latent cache per layer group, with layer-specific decoders defining effective keys.


- **Training-free calibration** uses offline ridge reduced-rank regression to jointly fit shared projections and layer-specific decoders, enabling direct scoring of the shared cache via query-side decoder absorption without reconstructing historical keys.


- **On DeepSeek-V3.2 and GLM-5**, training-free LatentIndex improves head-wise attention-mass recall over IndexCache by up to 3.28 percentage points while maintaining RULER and LongBench performance close to native DSA. With four-layer sharing, it reduces logical indexer-cache storage by 61.1%.


- **Hierarchical selection (HS)** variant restricts followers to a shared candidate set proposed by the anchor, achieving 2.30–2.72× decode indexer speedups over DSA across 8K–128K contexts while retaining most of LatentIndex's recall.


- **Training-aware experiments** on DeepSeek-V2-Lite demonstrate that shared latent indexers can be learned, achieving higher warmup recall and generally lower sparse-training language-modeling loss than IndexCache, with LM loss close to native DSA.

.


## Introduction and Theoretical Foundation

Sparse attention reduces core-attention computation from quadratic to $\mathcal{O}(Lk)$ over sequence length $L$, but introduces two sources of indexing overhead: repeated selection computation (the indexer scans the full prefix at each layer) and per-layer key storage. Cross-layer redundancy offers an opportunity to reuse indexing information; IndexCache exploits this by letting Shared layers reuse an anchor layer's selected indices, but this constrains multiple layers to the same token set, preventing layer-specific adaptation.



The paper draws on the design principle of Multi-head Latent Attention (MLA), which shows that a compact latent representation can be shared across attention heads while supporting head-specific attention. The central research question is: *Can we exploit cross-layer redundancy while retaining layer-specific token selection?* LatentIndex answers this by sharing continuous representations rather than discrete selections, allowing each layer to maintain its own scoring and top-k selection while sharing a common latent cache.



The theoretical foundation rests on the observation that nearby layers' indexer keys exhibit substantial redundancy, making them amenable to low-rank joint factorization. The paper also leverages MLA's query-side decoder absorption technique, which enables direct scoring of shared latents without materializing per-layer keys during inference.



## Methodology

**Shared Latent Cache.** Each group of consecutive layers $\mathcal{G}_g$ constructs one shared representation per token from the anchor layer's hidden state. For token position $t$, the latent cache is computed as:

$$
\mathbf{z}_{t}^{(g),N} = \mathbf{P}_{g,N}^{\top} \mathbf{x}_{t}^{(a_{g})}}, \qquad \mathbf{z}_{t}^{(g),R} = \mathbf{Q}_{t} \mathbf{P}_{g,R}^{\top} \mathbf{x}_{t}^{(a_{g})}},\tag{1}
$$

where $\mathbf{P}_{g,N} \in \mathbb{R}^{d_{\mathrm{model}}} \times r_N}$ and $\mathbf{P}_{g,R} \in \mathbb{R}^{d_{\mathrm{model}}} \times r_R}$ are group-specific projections, $\mathbf{Q}_t$ applies rotary position encoding, and the two branches are concatenated:

$$
\mathbf{c}_{t}^{(g)} = \left[ \begin{array}{c} \mathbf{z}_{t}^{(g),R} \\ \mathbf{z}_{t}^{(g),N} \end{array} \right] \in \mathbb{R}^{d_C}.\tag{2}
$$


**Layer-specific Key Reconstruction.** Each layer $\ell \in \mathcal{G}_g$ uses its own decoder to map the shared latent to effective indexer keys:

$$
\widehat{\mathbf{k}}_{t}^{(\ell),R} = \mathbf{D}^{(\ell),R} \mathbf{z}_{t}^{(g),R}, \qquad \widehat{\mathbf{k}}_{t}^{(\ell),N} = \mathbf{D}^{(\ell),N} \mathbf{z}_{t}^{(g),N} + \mathbf{b}^{(\ell)},\tag{3}
$$

with a rotary-structure preservation constraint:

$$
\mathbf{R}_{t} \mathbf{D}^{(\ell),R} = \mathbf{D}^{(\ell),R} \mathbf{Q}_{t},\tag{4}
$$

which is enforced through complex linear combinations within each rotary frequency.

**Query-side Decoder Absorption.** Decoders are absorbed into queries to score the shared cache directly:

$$
\widetilde{\mathbf{q}}_{s,h}^{(\ell)} = \left(\mathbf{D}^{(\ell)}\right)^{\top} \mathbf{q}_{s,h}^{(\ell)} \in \mathbb{R}^{d_C}, \qquad \beta_{s,h}^{(\ell)} = \left(\mathbf{q}_{s,h}^{(\ell),N}\right)^{\top} \mathbf{b}^{(\ell)},\tag{6}
$$

yielding exact query–key dot products:

$$
\left(\mathbf{q}_{s,h}^{(\ell)}\right)^{\top} \widehat{\mathbf{k}}_{t}^{(\ell)} = \left(\widetilde{\mathbf{q}}_{s,h}^{(\ell)}\right)^{\top} \mathbf{c}_{t}^{(g)} + \beta_{s,h}^{(\ell)}.\tag{7}
$$


**Scoring and Selection.** Each layer independently scores the shared cache over valid causal positions $\mathcal{V}_s$:

$$
u_{s,t}^{(\ell)} = \sum_{h=1}^{H_I} w_{s,h}^{(\ell)} \mathrm{ReLU}\bigg(\left(\widetilde{\mathbf{q}}_{s,h}^{(\ell)}\right)^{\top} \mathbf{c}_{t}^{(g)} + \beta_{s,h}^{(\ell)} \bigg), \qquad t \in \mathcal{V}_s,\tag{8}
$$

and selects its own top-k set $\mathcal{S}_s^{(\ell)} = \operatorname{TopK}_{t \in \mathcal{V}_s}\left(u_{s,t}^{(\ell)}, \min(k, |\mathcal{V}_s|)\right)$.


**Training-free Calibration.** Offline calibration uses ridge reduced-rank regression to jointly fit a shared projection and layer-specific decoders, minimizing key-reconstruction error under latent-rank constraints. NoPE and RoPE branches are fit separately, with closed-form solutions requiring no gradient-based optimization.



**Hierarchical Selection.** The anchor scores all valid positions and forms a candidate set $\mathcal{C}_s^{(g)}$ from the top-$B$ positions (where $B \ge k$). Each follower independently rescores these candidates using its layer-specific scorer, reducing scored positions per query from $mL$ to $L + (m-1)\min(B, L)$ per group of $m$ layers.



**Training-aware Instantiation.** Training proceeds in two stages: (1) dense-attention warmup with frozen backbone, training indexers to match dense attention via KL divergence; (2) sparse training that updates both backbone and indexers using selected routes, with continued KL supervision over selected positions.



## Empirical Validation / Results

**Attention-mass Recall.** LatentIndex outperforms IndexCache at every evaluated length on both models, with gains increasing toward longer contexts:

| **Method** | **4K** | **8K** | **16K** | **32K** | **64K** | **128K** |
|---|---|---|---|---|---|---|
| **DeepSeek-V3.2** | | | | | | |
| DSA | 99.33 | 97.40 | 95.09 | 93.48 | 91.10 | 90.00 |
| IndexCache | 98.86 | 95.86 | 92.49 | 90.26 | 86.68 | 85.06 |
| LatentIndex | 99.21 | 96.97 | 94..34 | 92..40 | | 89..61 | | 88..34 |
| $\Delta$ vs. IndexCache | +0...35 | +1..11 | +1..85 | +2..14 | +2..93 | +3..28 |

**Table 1:** Head-wise attention-mass recall on a common native DSA trajectory. At 128K, LatentIndex improves recall by  ​​3​​.​​28 percentage points on DeepSeek-V3.2 and 1​​.​​57 points on GLM-5.

.​


**RULER Performance.** LatentIndex closely preserves native DSA performance, with average scores only 0.05 and​0.​03 points lower on DeepSeek-V3.2 and GLM-5, respectively. It improves over IndexCache in aggregate on both models, with the largest gain on GLM-5 at 128K (96​​.​​81 vs.​ 90​​.​​15 for IndexCache​).


**LongBench Performance.** LatentIndex improves the average score over IndexCache by​ 0​​.​​73 points on DeepSeek-V3.2 and​ 0​​.​​30 points on GLM-5, while remaining close to native DSA. It improves all six categories relative to IndexCache on DeepSeek-V3.2, with the largest gain in few-shot learning.


**Cross-layer Sharing Analysis.** Within each four-layer group, LatentIndex reduces recall loss at every follower offset and context length. At 128K, average follower loss decreases by​ 73%, from​ 6​​.​​69 to​ 1​​.​​81 percentage points. Larger groups ($G = 8$) reduce recall for both methods, but LatentIndex with $G = 8$ achieves 86​​.​​21 recall, exceeding IndexCache with $G =  ​​8$ by​​ 4​​.​​07 points and even IndexCache with $G = 4$ by​​ 1​​.​​15 points.


**Hierarchical Selection Results.** With a fixed candidate budget of​ 8192:

| **Method** | **$R_{\text{head}} \uparrow$** | **RULER ↑** | **LongBench ↑** | **8K** | **16K** | **32K** | **64K** | **128K** |
|---|---|---|---|---|---|---|---|---|
| Native DSA | 94​​.​​40 |​ 96​​.​​54 |​ 55​​.​​24 |​ 48​​.​​70 |​ 52​​.​​06 |​ 48​​.​​28 |​ 48​​.​​43 |​ 49​​.​​60 |
| IndexCache |​ 91​​.​​54 |​ 96​​.​​40 |​ 54​​.​​61 |​ 14​​.​​31 |​ 13​​.​​29 |​ 14​​.​​35 |​ 14​​.​​28 |​ 13​​.​​40 |
| LatentIndex + HS |​ 93​​.​​19 |​ 96​​.​​58 |​ 54​​.​​81 |​ 19​​.​​89 (2​​.​​45×) |​ 19​​.​​13 (2​​.​​72×) |​ 20​​.​​55 (2​​.​​35×) |​ 21​​.​​10 (2​​.​​30×) |​ 20​​.​​48 (2​​.​​42×) |

**Table 4:** Hierarchical selection achieves 2​​.​​30–2​​.​​72× decode indexer speedups over native DSA while retaining most of LatentIndex's recall.


**Training-aware Results.** On DeepSeek-V2-Lite, LatentIndex achieves higher warmup Recall@2048 (approximately​ 1​​.​​3–1​​.​​5 percentage points below DSA, vs.​ 4​​.​​4–4​​.​​7 for IndexCache) and generally lower sparse-training LM loss than IndexCache, fluctuating closely around the DSA baseline.



## Theoretical and Practical Implications

**Storage Efficiency.** By storing one $d_C$-dimensional latent instead of $m$ separate $d_I$-dimensional keys per group, the payload ratio becomes $d_C/(m d_I)$. Including FP8 scales, the DeepSeek-V3.2 configuration reduces logical indexer-cache storage by​ 61​​.​​1%, addressing a key bottleneck in sparse attention serving.



**Selection Quality vs. Computation Trade-off.** LatentIndex demonstrates that sharing continuous representations rather than discrete selections enables efficient reuse while preserving layer-specific token selection. This insight extends MLA's latent-sharing principle from the head dimension to the layer dimension of indexers, offering a new perspective on cross-layer indexing.



**Practical Deployment.** The training-free calibration requires no gradient-based optimization, making it directly applicable to frozen pretrained models. Hierarchical selection provides a tunable quality–efficiency knob, where the candidate budget $B$ controls the trade-off between scoring cost and selection quality. The training-aware instantiation demonstrates that shared latent indexers can be learned from scratch, opening avenues for future model pretraining with integrated sparse attention.



## Conclusion

LatentIndex enables sparse-attention indexers to share continuous key representations across layers while retaining layer-specific token selection. Training-free results on DeepSeek-V3.2and GLM-5 show higher attention-mass recall than index reuse and downstream performance close to native DSA, with a 61​​.​​1% reduction in logical indexer-cache storage. Hierarchical selection reduces repeated full-prefix scoring through shared candidates, achieving decode indexer speedups while retaining layer-specific refinement. Training experiments provide complementary evidence that shared latent indexers can be learned while maintaining language-model loss close to DSA.



**Future directions** include: exploring larger sharing groups with higher latent dimensions, extending LatentIndex to other sparse-attention architectures, integrating the training-aware instantiation into large-scale pretraining pipelines, and investigating adaptive candidate budgets for hierarchical selection based on query or context characteristics.

---

_Markdown view of https://picx.dev/p/VoMFGs, served by PicX — AI-generated visual whiteboard summaries of research papers._
