Summary (Overview)

  • LatentIndex extends the latent-sharing principle of Multi-head Latent Attention (MLA) to sparse-attention indexers, enabling cross-layer sharing of continuous key representations while preserving layer-specific token selection. . The key innovation is replacing per-layer key caches with a shared latent cache per layer group, with layer-specific decoders defining effective keys.

  • Training-free calibration uses offline ridge reduced-rank regression to jointly fit shared projections and layer-specific decoders, enabling direct scoring of the shared cache via query-side decoder absorption without reconstructing historical keys.

  • On DeepSeek-V3.2 and GLM-5, training-free LatentIndex improves head-wise attention-mass recall over IndexCache by up to 3.28 percentage points while maintaining RULER and LongBench performance close to native DSA. With four-layer sharing, it reduces logical indexer-cache storage by 61.1%.

  • Hierarchical selection (HS) variant restricts followers to a shared candidate set proposed by the anchor, achieving 2.30–2.72× decode indexer speedups over DSA across 8K–128K contexts while retaining most of LatentIndex's recall.

  • Training-aware experiments on DeepSeek-V2-Lite demonstrate that shared latent indexers can be learned, achieving higher warmup recall and generally lower sparse-training language-modeling loss than IndexCache, with LM loss close to native DSA.

.

Introduction and Theoretical Foundation

Sparse attention reduces core-attention computation from quadratic to O(Lk)\mathcal{O}(Lk) over sequence length LL, but introduces two sources of indexing overhead: repeated selection computation (the indexer scans the full prefix at each layer) and per-layer key storage. Cross-layer redundancy offers an opportunity to reuse indexing information; IndexCache exploits this by letting Shared layers reuse an anchor layer's selected indices, but this constrains multiple layers to the same token set, preventing layer-specific adaptation.

The paper draws on the design principle of Multi-head Latent Attention (MLA), which shows that a compact latent representation can be shared across attention heads while supporting head-specific attention. The central research question is: Can we exploit cross-layer redundancy while retaining layer-specific token selection? LatentIndex answers this by sharing continuous representations rather than discrete selections, allowing each layer to maintain its own scoring and top-k selection while sharing a common latent cache.

The theoretical foundation rests on the observation that nearby layers' indexer keys exhibit substantial redundancy, making them amenable to low-rank joint factorization. The paper also leverages MLA's query-side decoder absorption technique, which enables direct scoring of shared latents without materializing per-layer keys during inference.

Methodology

Shared Latent Cache. Each group of consecutive layers Gg\mathcal{G}_g constructs one shared representation per token from the anchor layer's hidden state. For token position tt, the latent cache is computed as:

\mathbf{z}_{t}^{(g),N} = \mathbf{P}_{g,N}^{\top} \mathbf{x}_{t}^{(a_{g})}}, \qquad \mathbf{z}_{t}^{(g),R} = \mathbf{Q}_{t} \mathbf{P}_{g,R}^{\top} \mathbf{x}_{t}^{(a_{g})}},\tag{1}

where \mathbf{P}_{g,N} \in \mathbb{R}^{d_{\mathrm{model}}} \times r_N} and \mathbf{P}_{g,R} \in \mathbb{R}^{d_{\mathrm{model}}} \times r_R} are group-specific projections, Qt\mathbf{Q}_t applies rotary position encoding, and the two branches are concatenated:

ct(g)=[zt(g),Rzt(g),N]∈RdC.(2)\mathbf{c}_{t}^{(g)} = \left[ \begin{array}{c} \mathbf{z}_{t}^{(g),R} \\ \mathbf{z}_{t}^{(g),N} \end{array} \right] \in \mathbb{R}^{d_C}.\tag{2}

Layer-specific Key Reconstruction. Each layer ℓ∈Gg\ell \in \mathcal{G}_g uses its own decoder to map the shared latent to effective indexer keys:

k^t(ℓ),R=D(ℓ),Rzt(g),R,k^t(ℓ),N=D(ℓ),Nzt(g),N+b(ℓ),(3)\widehat{\mathbf{k}}_{t}^{(\ell),R} = \mathbf{D}^{(\ell),R} \mathbf{z}_{t}^{(g),R}, \qquad \widehat{\mathbf{k}}_{t}^{(\ell),N} = \mathbf{D}^{(\ell),N} \mathbf{z}_{t}^{(g),N} + \mathbf{b}^{(\ell)},\tag{3}

with a rotary-structure preservation constraint:

RtD(ℓ),R=D(ℓ),RQt,(4)\mathbf{R}_{t} \mathbf{D}^{(\ell),R} = \mathbf{D}^{(\ell),R} \mathbf{Q}_{t},\tag{4}

which is enforced through complex linear combinations within each rotary frequency.

Query-side Decoder Absorption. Decoders are absorbed into queries to score the shared cache directly:

q~s,h(ℓ)=(D(ℓ))⊤qs,h(ℓ)∈RdC,βs,h(ℓ)=(qs,h(ℓ),N)⊤b(ℓ),(6)\widetilde{\mathbf{q}}_{s,h}^{(\ell)} = \left(\mathbf{D}^{(\ell)}\right)^{\top} \mathbf{q}_{s,h}^{(\ell)} \in \mathbb{R}^{d_C}, \qquad \beta_{s,h}^{(\ell)} = \left(\mathbf{q}_{s,h}^{(\ell),N}\right)^{\top} \mathbf{b}^{(\ell)},\tag{6}

yielding exact query–key dot products:

(qs,h(ℓ))⊤k^t(ℓ)=(q~s,h(ℓ))⊤ct(g)+βs,h(ℓ).(7)\left(\mathbf{q}_{s,h}^{(\ell)}\right)^{\top} \widehat{\mathbf{k}}_{t}^{(\ell)} = \left(\widetilde{\mathbf{q}}_{s,h}^{(\ell)}\right)^{\top} \mathbf{c}_{t}^{(g)} + \beta_{s,h}^{(\ell)}.\tag{7}

Scoring and Selection. Each layer independently scores the shared cache over valid causal positions Vs\mathcal{V}_s:

us,t(ℓ)=∑h=1HIws,h(ℓ)ReLU((q~s,h(ℓ))⊤ct(g)+βs,h(ℓ)),t∈Vs,(8)u_{s,t}^{(\ell)} = \sum_{h=1}^{H_I} w_{s,h}^{(\ell)} \mathrm{ReLU}\bigg(\left(\widetilde{\mathbf{q}}_{s,h}^{(\ell)}\right)^{\top} \mathbf{c}_{t}^{(g)} + \beta_{s,h}^{(\ell)} \bigg), \qquad t \in \mathcal{V}_s,\tag{8}

and selects its own top-k set Ss(ℓ)=TopK⁡t∈Vs(us,t(ℓ),min⁡(k,∣Vs∣))\mathcal{S}_s^{(\ell)} = \operatorname{TopK}_{t \in \mathcal{V}_s}\left(u_{s,t}^{(\ell)}, \min(k, |\mathcal{V}_s|)\right).

Training-free Calibration. Offline calibration uses ridge reduced-rank regression to jointly fit a shared projection and layer-specific decoders, minimizing key-reconstruction error under latent-rank constraints. NoPE and RoPE branches are fit separately, with closed-form solutions requiring no gradient-based optimization.

Hierarchical Selection. The anchor scores all valid positions and forms a candidate set Cs(g)\mathcal{C}_s^{(g)} from the top-BB positions (where B≥kB \ge k). Each follower independently rescores these candidates using its layer-specific scorer, reducing scored positions per query from mLmL to L+(m−1)min⁡(B,L)L + (m-1)\min(B, L) per group of mm layers.

Training-aware Instantiation. Training proceeds in two stages: (1) dense-attention warmup with frozen backbone, training indexers to match dense attention via KL divergence; (2) sparse training that updates both backbone and indexers using selected routes, with continued KL supervision over selected positions.

Empirical Validation / Results

Attention-mass Recall. LatentIndex outperforms IndexCache at every evaluated length on both models, with gains increasing toward longer contexts:

Method4K8K16K32K64K128K
DeepSeek-V3.2
DSA99.3397.4095.0993.4891.1090.00
IndexCache98.8695.8692.4990.2686.6885.06
LatentIndex99.2196.9794..3492..4089..61
Δ\Delta vs. IndexCache+0...35+1..11+1..85+2..14+2..93+3..28

Table 1: Head-wise attention-mass recall on a common native DSA trajectory. At 128K, LatentIndex improves recall by ​​3​​.​​28 percentage points on DeepSeek-V3.2 and 1​​.​​57 points on GLM-5.

.​

RULER Performance. LatentIndex closely preserves native DSA performance, with average scores only 0.05 and​0.​03 points lower on DeepSeek-V3.2 and GLM-5, respectively. It improves over IndexCache in aggregate on both models, with the largest gain on GLM-5 at 128K (96​​.​​81 vs.​ 90​​.​​15 for IndexCache​).

LongBench Performance. LatentIndex improves the average score over IndexCache by​ 0​​.​​73 points on DeepSeek-V3.2 and​ 0​​.​​30 points on GLM-5, while remaining close to native DSA. It improves all six categories relative to IndexCache on DeepSeek-V3.2, with the largest gain in few-shot learning.

Cross-layer Sharing Analysis. Within each four-layer group, LatentIndex reduces recall loss at every follower offset and context length. At 128K, average follower loss decreases by​ 73%, from​ 6​​.​​69 to​ 1​​.​​81 percentage points. Larger groups (G=8G = 8) reduce recall for both methods, but LatentIndex with G=8G = 8 achieves 86​​.​​21 recall, exceeding IndexCache with G=​​8G = ​​8 by​​ 4​​.​​07 points and even IndexCache with G=4G = 4 by​​ 1​​.​​15 points.

Hierarchical Selection Results. With a fixed candidate budget of​ 8192:

MethodRhead↑R_{\text{head}} \uparrowRULER ↑LongBench ↑8K16K32K64K128K
Native DSA94​​.​​40​ 96​​.​​54​ 55​​.​​24​ 48​​.​​70​ 52​​.​​06​ 48​​.​​28​ 48​​.​​43​ 49​​.​​60
IndexCache​ 91​​.​​54​ 96​​.​​40​ 54​​.​​61​ 14​​.​​31​ 13​​.​​29​ 14​​.​​35​ 14​​.​​28​ 13​​.​​40
LatentIndex + HS​ 93​​.​​19​ 96​​.​​58​ 54​​.​​81​ 19​​.​​89 (2​​.​​45×)​ 19​​.​​13 (2​​.​​72×)​ 20​​.​​55 (2​​.​​35×)​ 21​​.​​10 (2​​.​​30×)​ 20​​.​​48 (2​​.​​42×)

Table 4: Hierarchical selection achieves 2​​.​​30–2​​.​​72× decode indexer speedups over native DSA while retaining most of LatentIndex's recall.

Training-aware Results. On DeepSeek-V2-Lite, LatentIndex achieves higher warmup Recall@2048 (approximately​ 1​​.​​3–1​​.​​5 percentage points below DSA, vs.​ 4​​.​​4–4​​.​​7 for IndexCache) and generally lower sparse-training LM loss than IndexCache, fluctuating closely around the DSA baseline.

Theoretical and Practical Implications

Storage Efficiency. By storing one dCd_C-dimensional latent instead of mm separate dId_I-dimensional keys per group, the payload ratio becomes dC/(mdI)d_C/(m d_I). Including FP8 scales, the DeepSeek-V3.2 configuration reduces logical indexer-cache storage by​ 61​​.​​1%, addressing a key bottleneck in sparse attention serving.

Selection Quality vs. Computation Trade-off. LatentIndex demonstrates that sharing continuous representations rather than discrete selections enables efficient reuse while preserving layer-specific token selection. This insight extends MLA's latent-sharing principle from the head dimension to the layer dimension of indexers, offering a new perspective on cross-layer indexing.

Practical Deployment. The training-free calibration requires no gradient-based optimization, making it directly applicable to frozen pretrained models. Hierarchical selection provides a tunable quality–efficiency knob, where the candidate budget BB controls the trade-off between scoring cost and selection quality. The training-aware instantiation demonstrates that shared latent indexers can be learned from scratch, opening avenues for future model pretraining with integrated sparse attention.

Conclusion

LatentIndex enables sparse-attention indexers to share continuous key representations across layers while retaining layer-specific token selection. Training-free results on DeepSeek-V3.2and GLM-5 show higher attention-mass recall than index reuse and downstream performance close to native DSA, with a 61​​.​​1% reduction in logical indexer-cache storage. Hierarchical selection reduces repeated full-prefix scoring through shared candidates, achieving decode indexer speedups while retaining layer-specific refinement. Training experiments provide complementary evidence that shared latent indexers can be learned while maintaining language-model loss close to DSA.

Future directions include: exploring larger sharing groups with higher latent dimensions, extending LatentIndex to other sparse-attention architectures, integrating the training-aware instantiation into large-scale pretraining pipelines, and investigating adaptive candidate budgets for hierarchical selection based on query or context characteristics.

Related papers