Summary (Overview)
-
LatentIndex extends the latent-sharing principle of Multi-head Latent Attention (MLA) to sparse-attention indexers, enabling cross-layer sharing of continuous key representations while preserving layer-specific token selection. . The key innovation is replacing per-layer key caches with a shared latent cache per layer group, with layer-specific decoders defining effective keys.
-
Training-free calibration uses offline ridge reduced-rank regression to jointly fit shared projections and layer-specific decoders, enabling direct scoring of the shared cache via query-side decoder absorption without reconstructing historical keys.
-
On DeepSeek-V3.2 and GLM-5, training-free LatentIndex improves head-wise attention-mass recall over IndexCache by up to 3.28 percentage points while maintaining RULER and LongBench performance close to native DSA. With four-layer sharing, it reduces logical indexer-cache storage by 61.1%.
-
Hierarchical selection (HS) variant restricts followers to a shared candidate set proposed by the anchor, achieving 2.30–2.72× decode indexer speedups over DSA across 8K–128K contexts while retaining most of LatentIndex's recall.
-
Training-aware experiments on DeepSeek-V2-Lite demonstrate that shared latent indexers can be learned, achieving higher warmup recall and generally lower sparse-training language-modeling loss than IndexCache, with LM loss close to native DSA.
.
Introduction and Theoretical Foundation
Sparse attention reduces core-attention computation from quadratic to over sequence length , but introduces two sources of indexing overhead: repeated selection computation (the indexer scans the full prefix at each layer) and per-layer key storage. Cross-layer redundancy offers an opportunity to reuse indexing information; IndexCache exploits this by letting Shared layers reuse an anchor layer's selected indices, but this constrains multiple layers to the same token set, preventing layer-specific adaptation.
The paper draws on the design principle of Multi-head Latent Attention (MLA), which shows that a compact latent representation can be shared across attention heads while supporting head-specific attention. The central research question is: Can we exploit cross-layer redundancy while retaining layer-specific token selection? LatentIndex answers this by sharing continuous representations rather than discrete selections, allowing each layer to maintain its own scoring and top-k selection while sharing a common latent cache.
The theoretical foundation rests on the observation that nearby layers' indexer keys exhibit substantial redundancy, making them amenable to low-rank joint factorization. The paper also leverages MLA's query-side decoder absorption technique, which enables direct scoring of shared latents without materializing per-layer keys during inference.
Methodology
Shared Latent Cache. Each group of consecutive layers constructs one shared representation per token from the anchor layer's hidden state. For token position , the latent cache is computed as:
\mathbf{z}_{t}^{(g),N} = \mathbf{P}_{g,N}^{\top} \mathbf{x}_{t}^{(a_{g})}}, \qquad \mathbf{z}_{t}^{(g),R} = \mathbf{Q}_{t} \mathbf{P}_{g,R}^{\top} \mathbf{x}_{t}^{(a_{g})}},\tag{1}where \mathbf{P}_{g,N} \in \mathbb{R}^{d_{\mathrm{model}}} \times r_N} and \mathbf{P}_{g,R} \in \mathbb{R}^{d_{\mathrm{model}}} \times r_R} are group-specific projections, applies rotary position encoding, and the two branches are concatenated:
Layer-specific Key Reconstruction. Each layer uses its own decoder to map the shared latent to effective indexer keys:
with a rotary-structure preservation constraint:
which is enforced through complex linear combinations within each rotary frequency.
Query-side Decoder Absorption. Decoders are absorbed into queries to score the shared cache directly:
yielding exact query–key dot products:
Scoring and Selection. Each layer independently scores the shared cache over valid causal positions :
and selects its own top-k set .
Training-free Calibration. Offline calibration uses ridge reduced-rank regression to jointly fit a shared projection and layer-specific decoders, minimizing key-reconstruction error under latent-rank constraints. NoPE and RoPE branches are fit separately, with closed-form solutions requiring no gradient-based optimization.
Hierarchical Selection. The anchor scores all valid positions and forms a candidate set from the top- positions (where ). Each follower independently rescores these candidates using its layer-specific scorer, reducing scored positions per query from to per group of layers.
Training-aware Instantiation. Training proceeds in two stages: (1) dense-attention warmup with frozen backbone, training indexers to match dense attention via KL divergence; (2) sparse training that updates both backbone and indexers using selected routes, with continued KL supervision over selected positions.
Empirical Validation / Results
Attention-mass Recall. LatentIndex outperforms IndexCache at every evaluated length on both models, with gains increasing toward longer contexts:
| Method | 4K | 8K | 16K | 32K | 64K | 128K |
|---|---|---|---|---|---|---|
| DeepSeek-V3.2 | ||||||
| DSA | 99.33 | 97.40 | 95.09 | 93.48 | 91.10 | 90.00 |
| IndexCache | 98.86 | 95.86 | 92.49 | 90.26 | 86.68 | 85.06 |
| LatentIndex | 99.21 | 96.97 | 94..34 | 92..40 | 89..61 | |
| vs. IndexCache | +0...35 | +1..11 | +1..85 | +2..14 | +2..93 | +3..28 |
Table 1: Head-wise attention-mass recall on a common native DSA trajectory. At 128K, LatentIndex improves recall by 3.28 percentage points on DeepSeek-V3.2 and 1.57 points on GLM-5.
.
RULER Performance. LatentIndex closely preserves native DSA performance, with average scores only 0.05 and0.03 points lower on DeepSeek-V3.2 and GLM-5, respectively. It improves over IndexCache in aggregate on both models, with the largest gain on GLM-5 at 128K (96.81 vs. 90.15 for IndexCache).
LongBench Performance. LatentIndex improves the average score over IndexCache by 0.73 points on DeepSeek-V3.2 and 0.30 points on GLM-5, while remaining close to native DSA. It improves all six categories relative to IndexCache on DeepSeek-V3.2, with the largest gain in few-shot learning.
Cross-layer Sharing Analysis. Within each four-layer group, LatentIndex reduces recall loss at every follower offset and context length. At 128K, average follower loss decreases by 73%, from 6.69 to 1.81 percentage points. Larger groups () reduce recall for both methods, but LatentIndex with achieves 86.21 recall, exceeding IndexCache with by 4.07 points and even IndexCache with by 1.15 points.
Hierarchical Selection Results. With a fixed candidate budget of 8192:
| Method | RULER ↑ | LongBench ↑ | 8K | 16K | 32K | 64K | 128K | |
|---|---|---|---|---|---|---|---|---|
| Native DSA | 94.40 | 96.54 | 55.24 | 48.70 | 52.06 | 48.28 | 48.43 | 49.60 |
| IndexCache | 91.54 | 96.40 | 54.61 | 14.31 | 13.29 | 14.35 | 14.28 | 13.40 |
| LatentIndex + HS | 93.19 | 96.58 | 54.81 | 19.89 (2.45×) | 19.13 (2.72×) | 20.55 (2.35×) | 21.10 (2.30×) | 20.48 (2.42×) |
Table 4: Hierarchical selection achieves 2.30–2.72× decode indexer speedups over native DSA while retaining most of LatentIndex's recall.
Training-aware Results. On DeepSeek-V2-Lite, LatentIndex achieves higher warmup Recall@2048 (approximately 1.3–1.5 percentage points below DSA, vs. 4.4–4.7 for IndexCache) and generally lower sparse-training LM loss than IndexCache, fluctuating closely around the DSA baseline.
Theoretical and Practical Implications
Storage Efficiency. By storing one -dimensional latent instead of separate -dimensional keys per group, the payload ratio becomes . Including FP8 scales, the DeepSeek-V3.2 configuration reduces logical indexer-cache storage by 61.1%, addressing a key bottleneck in sparse attention serving.
Selection Quality vs. Computation Trade-off. LatentIndex demonstrates that sharing continuous representations rather than discrete selections enables efficient reuse while preserving layer-specific token selection. This insight extends MLA's latent-sharing principle from the head dimension to the layer dimension of indexers, offering a new perspective on cross-layer indexing.
Practical Deployment. The training-free calibration requires no gradient-based optimization, making it directly applicable to frozen pretrained models. Hierarchical selection provides a tunable quality–efficiency knob, where the candidate budget controls the trade-off between scoring cost and selection quality. The training-aware instantiation demonstrates that shared latent indexers can be learned from scratch, opening avenues for future model pretraining with integrated sparse attention.
Conclusion
LatentIndex enables sparse-attention indexers to share continuous key representations across layers while retaining layer-specific token selection. Training-free results on DeepSeek-V3.2and GLM-5 show higher attention-mass recall than index reuse and downstream performance close to native DSA, with a 61.1% reduction in logical indexer-cache storage. Hierarchical selection reduces repeated full-prefix scoring through shared candidates, achieving decode indexer speedups while retaining layer-specific refinement. Training experiments provide complementary evidence that shared latent indexers can be learned while maintaining language-model loss close to DSA.
Future directions include: exploring larger sharing groups with higher latent dimensions, extending LatentIndex to other sparse-attention architectures, integrating the training-aware instantiation into large-scale pretraining pipelines, and investigating adaptive candidate budgets for hierarchical selection based on query or context characteristics.
Related papers
- Harness Evolution Hits a Ceiling: When Weight Training Should Begin
Harness evolution fixes process failures like loops and blocked calls, while weight training fixes content failures, with gains transferring only when edits change what the model writes.
- TestJack: Should you trust the results in coding benchmarks? Agentic Coding Benchmarks Auditing via Evaluator Evolution
TESTJACK refutes 34.4% of benchmark-passing LLM coding trials by evolving tests from prompt requirements, cutting reported resolution rates from 50.6% to 33.2%.
- Stateless Language Agents: Scaling Long-Horizon Automated Research
Stateless Language Agents, where the harness owns all research state and reconstructs fresh contexts per invocation, outperform stateful agent frameworks on long-horizon tasks, reaching baseline final performance with over 84% fewer tokens.