What Does the Recurrent State Remember: A Memory Anatomy of Linear Attention
Highlights of This Issue
The strongest signal this issue is a set of mechanistic works simultaneously asking the same question: what exactly does the linear/recurrent state remember, and how do hybrid architectures encode position? How Linear Attention Remembers uses three types of causal interventions—donor-state swap, write blocking, and output patch—on pretrained GLA/GDN to separate the write, retention, and read phases of recurrent memory, and quantitatively proves that in hybrid architectures, the runtime memory directly supporting recall is almost entirely transferred to the full-attention KV state (KV recovery >99% vs. recurrent <1%)—directly advancing the channel-level separation of Issue 5's 'What Attention Recalls'. Anatomy of Associative Recall performs a single-knob decomposition under a fixed state budget, finding that short convolution is the dominant lever (+0.47/+0.44), the advantage of rank-1 transfer shrinks to +0.034 with convolution, and proves that the interference wall is a training coverage gap rather than a capacity issue—both works jointly point to the core concern of 'who actually carries the memory of the recurrent state'.
Mechanistic work on positional encoding is equally dense. How Local Mixing Encodes Relative Position proves that local mixing induces a recency bias in the residual stream that is preserved across sequence lengths (contrary to the length dilution of global NoPE), and is selected by global NoPE attention logits, providing a testable design insight that 'the smaller the window, the stronger the recency bias and the lower the validation loss'; Shifting Mechanisms uses counterfactual activation patching to prove that PE choice shifts the retrieval mechanism from position-dominated to semantics-dominated, revealing that SWA NoPE's long-context gains are a retrieval strategy shift rather than a uniform improvement—actually degrading on competing key discrimination.
Two new axes emerge on the compression side. KV-Kaizen is the first to unify three compression axes—depth/rank/precision—into a budget-conditioned learnable layer-wise selector, without evicting tokens and composable multiplicatively with eviction methods; SuffixReplay uses sparse anchors plus suffix replay to achieve checkpoint-free page-wise prefix reuse, directly advancing Issue 6's Tail-Replay. Counterexamples to Local Reconstruction Gain provides an important negative result: local reconstruction gain cannot serve as a proxy for final fidelity—a warning at the evaluation protocol level for all sparse selection and eviction methods targeting local reconstruction.
Community and Developments
A direct operation-level fix appears on the architecture side for sink avoidance. Softpick uses a rectified softmax with the sum-to-one constraint removed to explicitly prevent sink formation, achieving 0% sink rate on 340M/1.8B and improving low-bit quantization robustness—contrasting with Issue 7's Elastic Threshold Attention uniform floor claim, but note that at 1.8B scale performance does not reach softmax and long-context retrieval is not improved.
Two cautionary reviews on long-context evaluation validity are worth attention. Is Half Your Context Window Still Marketing in 2026? points out that usable context (by NoLiMa 85% retention) decays much earlier than advertised windows; Half the Cache, Twice the Context gives a concrete negative result: FP8 storage + BF16 accumulation KV quantization at 128K drops long-range retrieval from 91% to 13% without surfacing errors—standard perplexity and short-context metrics completely miss such catastrophic failures, consistent with this issue's reservations about 'lossless quantization' conclusions.
Open Questions
- How Linear Attention Remembers proves that in hybrid architectures, recall is directly carried by full-attention KV, and the recurrent state carries almost nothing directly—if full-attention layers exist, are recurrent layers redundant for retrieval? How does this unify with Issue 5's 'What Attention Recalls' channel-level separation into inter-layer ratio design?
- Anatomy of Associative Recall's interference wall is a training coverage gap rather than a capacity issue—does this conclusion hold beyond toy scale? Can distance curricula fix long-context retrieval degradation at real scale?
- Counterexamples to Local Reconstruction Gain: if local reconstruction gain is not a valid proxy for final fidelity, what objective should block-sparse selection and KV eviction optimize? Can Issue 3's 'Trust the Mass' selection ceiling framework provide an alternative metric?
- KV-Kaizen's budget-conditioned selector versus Issue 6's DeepSeek-V4.1-Flash static pattern sharing and Issue 7's HySparse2 two-level sharing—can 'compression axes as learnable decisions' be jointly optimized with cross-layer KV sharing, or do they conflict in gradients?
Papers in this issue
Linear attention memory stores facts via concentrated content-specific writes and retrieves them via focused query-time reads, with hybrid models shifting recall to full-attention KV caches.
Editor's noteUses three types of causal interventions—donor-state swap, write blocking, and output patch—on pretrained GLA/GDN to separate the write, retention, and read phases of recurrent memory, and quantitatively proves that in hybrid architectures, the runtime memory directly supporting recall is almost entirely transferred to the full-attention KV state (KV recovery >99% vs. recurrent <1%). Relative to Issue 5's 'What Attention Recalls' channel-level intervention, the increment is pushing interventions to the single-head write/read lifecycle and providing quantitative evidence of memory substrate transfer—a must-read for anyone doing hybrid layer ratio design.
At matched state, recall in recurrent cells is dominated by a short convolution, while an apparent interference wall is a training-coverage gap fixable by curriculum.
Editor's noteFirst single-knob decomposition of associative memory in linear attention/SSM under a fixed state budget: short convolution is the dominant lever (+0.47/+0.44), the advantage of rank-1 transfer shrinks to +0.034 with convolution, and state-matched Mamba-2 can match rank-1 units, thereby refuting architectural category claims. More critically, it finds that the interference wall (haystack retrieval at random at the shortest length) is a training coverage gap rather than a capacity issue, and a distance curriculum can improve the unmodified architecture from 0.021 to 1.000. Relative to Issue 5's 'What Attention Recalls' channel-level separation, the increment is providing factor decomposition under fixed state and causal evidence from the training side.
Hybrid architectures with local mixing layers and NoPE global attention implicitly learn relative position encodings via recency bias, enabling superior length extrapolation.
Editor's noteProvides a mechanistic explanation of how hybrid architectures (SWA + global NoPE) implicitly encode relative position: local mixing induces a recency bias in the residual stream that is preserved across sequence lengths (contrary to the length dilution of global NoPE), is selected by global NoPE attention logits, and strengthens with depth. The key design insight is that the smaller the window, the stronger the recency bias and the lower the validation loss—providing a testable mechanistic basis for window hyperparameter selection, and the mechanism holds on 4096+ tokens.
KV-Kaizen learns per-layer cache compression choices across depth, rank, and precision, achieving 4x compression with no accuracy loss on models 7B and larger.
Editor's noteFirst unification of three compression axes—depth (cross-layer sharing), rank (MLA truncation), precision (quantization)—into a learnable layer-wise selector, jointly optimized under a single bit-per-token budget, and supporting budget conditioning (one selector serving multiple compression ratios). Relative to Issue 6's DeepSeek-V4.1-Flash static pattern sharing and Issue 7's HySparse2 two-level sharing, the increment is making the choice of compression axes itself a context-adaptive learnable decision, without evicting tokens and composable multiplicatively with eviction methods. Provides compute-matched comparisons at 7B/14B and RULER 16K degradation curves.
Positional encoding choice determines whether language models retrieve positionally or semantically, with SWA NoPE shifting to semantic retrieval and trading long-context gains for weaker competing-key discrimination.
Editor's noteUses counterfactual activation patching for mechanism attribution (positional/lexical/reflexive) to prove that PE choice systematically changes the retrieval mechanism the model relies on: RoPE models are primarily positional, PE-mixed models are primarily semantic, and reveals SWA NoPE's long-context gains as a retrieval strategy shift rather than a uniform improvement—actually degrading on competing key discrimination. Relative to Issue 1's 'Large Window Laziness' and Issue 4's 'Modern Transformers' frequency-axis dichotomy, the increment is pushing the impact of PE choice on retrieval to the mechanism attribution level, an efficient negative-result type.
Local attention-output reconstruction gains do not guarantee final-model fidelity, as residual completion can improve local error while worsening dense-model KL divergence.
Editor's noteSystematically proves via constructed counterexamples that local reconstruction gain cannot serve as a proxy for final fidelity of residual completion, directly challenging the local reconstruction objective commonly relied upon in block-sparse selection and KV eviction. Relative to Issue 7's CompKV joint optimization and Issue 3's 'Trust the Mass' selection ceiling, the increment is directly providing counterexamples to the validity of the local reconstruction objective—a warning at the evaluation protocol level for any method optimizing local reconstruction.
Triadic linear attention raises RNN memory state to a third-order tensor via a second key, matching Transformers at 64k context with minimal parameter overhead.
Editor's noteFirst generalization of linear attention's matrix state to a third-order tensor state: using the triadic outer product of a second key to increase state capacity from d to d·E, with parameter overhead of only two projections, and providing a chunkwise parallel form so that E-fold state growth does not incur E-fold computation. Relative to Issue 5's Kalman Delta Networks confidence state, the increment is making the state shape itself a scalable design axis. Provides compute-matched comparisons at 400M/1.3B and PG19 position-wise degradation curves.
SuffixReplay enables fine-grained prefix caching in hybrid LLMs by reconstructing linear-attention states from sparse anchors, cutting storage by 2x and TTFT by up to 70%.
Editor's noteFirst implementation of fine-grained page-boundary reuse for hybrid LLM prefix caching without materializing recurrent state checkpoints: stores layer-wise token-wise sparse anchors, and exploits the forgetting property of recurrent decay to reconstruct linear attention states with recent suffix replay. Relative to Issue 6's Tail-Replay exact FA KV cache + short recent suffix replay, the increment is introducing sparse anchors and achieving checkpoint-free reuse at each page boundary, with compute-matched comparisons against SGLang native caching and LongBench/RULER degradation curves.
On-demand attention uses a lightweight recall head to predict when global attention helps, recovering most quality with up to 2.65x decoding throughput.
Editor's noteFirst to turn 'whether to invoke global attention' into a learnable decision regressed from the frozen pretrained model's local decoding state, supervised by paired Full-Local NLL gain and retaining full KV history. Relative to Issue 7's 'Language Models Can Control Their Own Attention' explicit text protocol, the increment is that the selection signal comes from hidden states after local computation rather than model declarations, training requires only a lightweight recall head with frozen backbone, and provides actual decoding speedup with GPU-side conditional execution in vLLM and RULER 4K-64K degradation curves.








