Full text not available for this paper

Summary (Overview)

  • Core finding: This paper provides concrete counterexamples showing that improving local attention-output reconstruction in residual completion methods does not necessarily improve final model fidelity (KL divergence to the dense FULL model).
  • Methodology: The authors study training-free RESA and learned Top-K+φ residual completion estimators on frozen backbone language models (Llama-3.2, Qwen3), using single-layer interventions at 64K-token contexts with K=16.
  • Key result: Two prespecified interventions (ϕ L15 and RESA L23, both Qwen3-0.6B/Multi-LexSum) show positive local reconstruction gain but negative final KL fidelity relative to the all-abstain EXACT TOP-K baseline, on both discovery and prompt-token-disjoint holdout requests.
  • Control experiment: Exact restoration (dense attention) at the same layer improves final fidelity, showing the reversal is specific to approximate completion, not the layer itself.
  • Diagnostic utility: A task-independent negative-G masking rule repairs completion models in 11/12 fidelity comparisons, but repaired models do not consistently outperform EXACT TOP-K, showing local reconstruction is informative but not a substitute for final-fidelity measurement.

Introduction and Theoretical Foundation

Background: Query-aware sparse attention reduces inference cost by evaluating only a small, query-dependent subset of stored keys/values (e.g., QUEST, SparQ). Residual completion methods (e.g., RESA, Top-K+φ) estimate the aggregate attention contribution of tokens omitted from the exact sparse computation and add this estimate to the sparse output.

Core Question: Does improving a layer's attention-output reconstruction on the same incoming Q/K/V and selected support necessarily improve final model output fidelity?

Key distinction:

  • Local reconstruction: Whether completion makes attention output closer to dense attention at the selected layer on measured decode steps
  • Final fidelity: Whether the model's next-token distribution is closer to the dense FULL model, measured via KL divergence

Theoretical motivation: After the selected layer, residual blocks and attention layers transform the intervention and may change subsequent query-dependent token selections. A correction can therefore reduce same-input attention-output error while increasing divergence at the final model output. Prior work (Delta Attention, RippleKV) motivates examining downstream consequences of sparse-attention perturbations.

Definitions:

The local reconstruction gain is defined as:

g=1−∥Odense−OC∥2∥Odense−OTK∥2g = 1 - \frac{\|O_{\text{dense}} - O_{C}\|_2}{\|O_{\text{dense}} - O_{TK}\|_2}

where OdenseO_{\text{dense}}, OTKO_{TK}, and OCO_{C} denote the post-WOW_O dense attention output, Exact Top-K attention output, and attention output with residual completion, respectively, computed from the same incoming Q/K/V. g>0g > 0 exactly when completion is closer to dense attention than EXACT TOP-K under squared reconstruction error.

The final-fidelity gain is:

Fi=1∣Ti∣∑t∈Ti[DKL(pitFull∥pitTK)−DKL(pitFull∥pitC)]F_i = \frac{1}{|T_i|} \sum_{t \in T_i} \left[ D_{KL}(p^{\text{Full}}_{it} \| p^{TK}_{it}) - D_{KL}(p^{\text{Full}}_{it} \| p^{C}_{it}) \right]

where pTKp^{TK} is the all-abstain EXACT TOP-K path and pCp^C is the completion path. Positive FF favors completion; negative FF favors abstention.

Methodology

Models and estimators:

  • Backbones: Llama-3.2-1B/3B-Instruct and Qwen3-0.6B/1.7B (all frozen)
  • Estimators: Top-K+φ (learned positive-feature estimator) and RESA (training-free prior)
  • Tasks: RULER NIAH Multikey-2, FWE, HELMET Multi-LexSum (selected based only on FULL–EXACT TOP-K gap)
  • Setting: 65,536-token prompt, K=16 (84 exact prompt tokens: 4 sink + 64 recent + 16 Top-K)

Screening protocol:

  • 36 single-layer actions (3 layers × 3 tasks × 2 models × 2 estimators), each on 50 requests
  • Two actions survived the 72-direction screening multiplicity correction (α/72, α=0.05): ϕ L15 and RESA L23, both Qwen3-0.6B/Multi-LexSum
  • Action identities fixed before direct-runtime follow-up

Direct-runtime measurement:

  • FULL, single-layer completion intervention, and all-abstain EXACT TOP-K executed with the same estimator-specific sparse implementation
  • Local measurement samples up to 32 prediction steps at predetermined evenly spaced positions
  • At measured steps, EXACT TOP-K and completion have bitwise-equal incoming hidden states, Q/K/V, and selected token IDs
  • Outputs actually returned by attention modules are observed (not recomputed offline)
  • Non-interference check: first two requests compared with measurement disabled

Exact-restoration control: Replaces the selected layer's sparse attention with dense attention, leaving other layers sparse. Uses block-state metric GblockG_{\text{block}} comparing hidden-state error to the global FULL trajectory.

Task-independent masking: For each model, 30 unlabeled 64K calibration sequences drawn from FineWeb, arXiv Summarization, and BIGPATENT. Freeze layers satisfying:

Ggeneric,ℓ<0G_{\text{generic},\ell} < 0

where Ggeneric,ℓG_{\text{generic},\ell} is the mean post-WOW_O local reconstruction gain for layer ℓ\ell. Completion is disabled at these layers.

Statistical methods: Paired resampling of requests; 100,000 bootstrap draws; four-direction family per split with cutoff α/4; deterministic bootstrap streams.

Empirical Validation / Results

Direct-runtime local improvement and final degradation

Table 1: Local reconstruction gain and final-fidelity gain measured in the same execution (Qwen3-0.6B/Multi-LexSum, 64K/K=16, 50 requests per split):

SplitActionLocal reconstruction G ↑Final-fidelity F ↑Requests G_i > 0, F_i < 0
Discoveryϕ L150.383 [0.365, 0.401]−0.0153 [−0.0213, −0.0098]41/50
DiscoveryRESA L230.104 [0.087, 0.121]−0.0029 [−0.0037, −0.0020]43/50
Holdoutϕ L150.345 [0.244, 0.404]−0.0222 [−0.0286, −0.0160]43/50
HoldoutRESA L230.097 [0.072, 0.119]−0.0027 [−0.0036, −0.0018]38/50

All four action–split comparisons retain the local-positive/final-negative direction under the four-direction sensitivity analysis. The joint counts show the same ordering occurs within individual request summaries.

Exact restoration vs. approximate completion

Table 2: Approximate completion versus exact restoration (separate discovery-set controls):

ActionInterventionG_block ↑Final-fidelity F ↑
ϕ L15Completion−0.176 [−0.194, −0.158]−0.0153 [−0.0213, −0.0097]
ϕ L15Exact restoration0.067 [0.063, 0.072]0.0157 [0.0111, 0.0206]
RESA L23Completion−0.018 [−0.023, −0.014]−0.0029 [−0.0037, −0.0020]
RESA L23Exact restoration0.043 [0.037, 0.049]0.0104 [0.0085, 0.0124]

Exact restoration reduces final KL by ~4.7% (ϕ L15) and ~3.1% (RESA L23), while approximate completion increases it by ~4.6% and ~0.9%, respectively.

Repairing completion does not establish advantage over EXACT TOP-K

Table 3: Repair and comparison with EXACT TOP-K:

Model / est.OffRepair FRepair Uvs. TK Fvs. TK U
Qwen3-0.6B / ϕ92/1/01/2/02/1/01/2/0
Qwen3-0.6B / RESA123/0/02/1/00/1/20/2/1
Qwen3-1.7B / ϕ53/0/02/1/01/1/10/2/1
Qwen3-1.7B / RESA123/0/03/0/00/1/20/3/0

Counts are positive/unresolved/negative paired 95% intervals over three tasks.

Key example: Qwen3-0.6B/RESA on Multi-LexSum gains 0.733 in final fidelity relative to unmodified completion, but the repaired hybrid is still 0.076 worse than EXACT TOP-K.

Alternative local metrics: Row-mean G is CI-positive in only 2/4 direct comparisons, energy-pooled G in 3/4, and absolute local error reduction in 3/4—all three are unresolved in the ϕ holdout. The central result is specific to the prespecified within-request-median criterion.

Theoretical and Practical Implications

Theoretical significance:

  • Provides the first concrete counterexamples showing that local reconstruction gain is not a valid proxy for final-model fidelity in residual completion
  • The exact-restoration control shows the reversal is specific to approximate completion—moving a layer toward dense attention is not inherently harmful
  • Local reconstruction remains useful as a diagnostic (repairs models in 11/12 cases) but does not determine final-model ordering
  • The negative-G masking rule is estimator-dependent: it transfers poorly to PISA-0TH (only 2 positive, 5 negative fidelity repairs)

Practical implications:

  • Evaluation of residual completion methods should include final-model fidelity measurement, not just local attention-output error
  • Layerwise local reconstruction diagnostics can identify failure layers for repair, but repaired models may still underperform simple abstention (EXACT TOP-K)
  • The mismatch is condition-dependent: all 12 Llama conditions show CI-positive local and final fidelity gains, while 8 Qwen conditions combine CI-positive local gain with CI-negative final fidelity

Limitations:

  • Direct-runtime core contains only two selected Qwen3-0.6B/Multi-LexSum interventions at 64K/K=16
  • K ∈ {8, 16, 32} covers nearby aggressive budgets, not substantially denser support
  • The central result is specific to the prespecified within-request-median criterion; alternative local weightings are less uniform
  • No targeted one-layer utility loss is statistically resolved; later support can change after intervention

Conclusion

The paper establishes that local attention reconstruction quality and final-model fidelity are related but not interchangeable endpoints. A completion method can improve one without improving the other. Specifically:

  1. For two fixed residual-completion interventions, positive prespecified local reconstruction gain coexists with worse final dense-model fidelity on both discovery and holdout requests.
  2. Exact restoration at the same layer produces the opposite final effect, showing the reversal is specific to approximate completion.
  3. Negative-G masking often repairs broader completion models without guaranteeing an advantage over EXACT TOP-K.

The authors conclude that local reconstruction is useful evidence about an estimator, but it does not by itself predict the fidelity of the final model output. Future work should focus on understanding the downstream transformations that produce this divergence and developing evaluation protocols that measure both local and final endpoints.

Related papers