Full text not available for this paper

Summary (Overview)

  • Core Question: How does fixed-size recurrent state in linear attention architectures function as a memory system for storing, retaining, and retrieving factual information?
  • Key Finding: Linear attention memory operates through concentrated, content-specific write pathways at fact time and concentrated query-time read pathways at recall, rather than diffuse or token-addressable storage.
  • Selective Access with Coupling: Multiple facts can remain selectively accessible within a shared recurrent state, but exhibit cross-fact causal coupling rather than independent KV-like storage.
  • Capacity Limits: Recall and targeted editability degrade with increasing memory load, while elapsed context alone has substantially smaller effects; interference is shaped by overlap between subsequent writes.
  • Hybrid Architecture Shift: In hybrid models combining recurrent layers with full attention, the runtime memory substrate directly supporting recall shifts predominantly to the full-attention KV cache.

Introduction and Theoretical Foundation

The paper addresses a fundamental tension in efficient language model design. Standard attention retains key-value (KV) representations for every preceding token, causing memory to grow linearly with sequence length—a critical bottleneck for long-context inference in both large-scale serving and resource-constrained settings like mobile devices and robotics.

Linear attention offers an alternative: instead of retaining per-token KV representations, it summarizes past information in a fixed-size recurrent state whose size does not grow with sequence length. However, this efficiency fundamentally changes how information is represented:

"In a conventional KV cache, representations associated with individual tokens remain explicitly available to subsequent attention operations. In recurrent linear attention, information from many past tokens is continually updated and accumulated within the same finite state."

The paper poses several key questions:

  • Which state updates carry remembered content?
  • How is that information accessed at recall time?
  • How do multiple facts coexist within a shared state?
  • What causes retention to degrade as more information accumulates?
  • Does recurrent state remain the primary memory substrate when token-addressable full attention is available?

The theoretical foundation builds on prior work connecting linear attention to fast-weight programmers and associative memory interpretations (Schlag et al., 2021; Irie et al., 2021), with architectural improvements through data-dependent gating (GLA), delta-rule updates (DeltaNet), and their combination (Gated DeltaNet/GDN).

Methodology

Analytical Framework

The paper develops an analytical decomposition of recurrent state memory. For a recurrent attention head, the state update is:

St=Tt(St−1)+Wt,St∈Rdk×dvS_t = T_t(S_{t-1}) + W_t, \quad S_t \in \mathbb{R}^{d_k \times d_v}

where TtT_t is linear in the previous state (e.g., Tt(S)=Gt⊙ST_t(S) = G_t \odot S for multiplicative gating), and a common rank-one write takes the form Wt=utvt⊤W_t = u_t v_t^\top.

The read operation is:

ot=qt⊤St,qt∈Rdk,ot∈Rdvo_t = q_t^\top S_t, \quad q_t \in \mathbb{R}^{d_k}, \quad o_t \in \mathbb{R}^{d_v}

The state can be decomposed into contributions from earlier positions:

St=∑i=1tCi(t)S_t = \sum_{i=1}^{t} C_i^{(t)}

where Ci(t)=(Tt∘Tt−1∘⋯∘Ti+1)(Ci(i))C_i^{(t)} = (T_t \circ T_{t-1} \circ \cdots \circ T_{i+1})(C_i^{(i)}) and Ci(i)=WiC_i^{(i)} = W_i.

Key theoretical insights from this decomposition:

  1. Past information shares a common state—contributions from different positions coexist without explicit token indexing
  2. Storage and access are distinct—a contribution can remain in StS_t without affecting current output
  3. The aggregate state does not uniquely reveal source contributions—the mapping from individual contributions to their sum is many-to-one

Experimental Design

Models Evaluated:

  • GLA-340M (Gated Linear Attention, 341.7M params)
  • GDN-340M (Gated DeltaNet, 399.5M params)
  • GDN-1.3B (1.466B params)
  • Hybrid GDN-340M (near-matched, 8 of 24 blocks replaced with full attention)
  • Qwen3.5-4B and Qwen3.5-9B (larger hybrids)

Task: Controlled associative recall using WikiText-103 passages with key–value associations embedded in natural text, evaluated via four-choice recall accuracy and target logit margin.

Causal Intervention Framework (summarized in Figure 1 of the paper):

  1. Fact-time interventions: Block or replace the recipient's fact-time write with donor counterpart
  2. State evolution: Allow state to evolve under subsequent tokens
  3. Pre-query state interventions: Swap recipient's pre-query state with donor state
  4. Query-time read interventions: Patch or restore the selected head's output
  5. Multi-fact extension: Test selective recall with multiple facts sharing recurrent state

Head Selection: Development-set screening of all recurrent heads (96/96/192 across models) using donor-state swaps, measured by the score:

Jℓ,h=116∑i=116[(zi,diℓ,h−zi,riℓ,h)⋅(zi,diclean−zi,riclean)]iJ_{\ell,h} = \frac{1}{16}\sum_{i=1}^{16}\left[(z^{\ell,h}_{i,d_i} - z^{\ell,h}_{i,r_i}) \cdot (z^{\text{clean}}_{i,d_i} - z^{\text{clean}}_{i,r_i})\right]_i

Selected heads: L21H3 (GLA-340M), L20H2 (GDN-340M), L20H3 (GDN-1.3B).

Empirical Validation / Results

1. Content-Specific Writes Persist in Recurrent State

  • Blocking the fact-time write reduces recall by 36–79 percentage points across models
  • Replacing the recipient's update with donor counterpart makes the donor value the answer in 68–90% of cases
  • Control interventions (norm-matched random, control-head updates) do not reproduce this effect
  • The selected head is part of an interacting recurrent circuit rather than an isolated fact store—blocking the hub head alone can impair recall more than blocking all recurrent heads together, with strongly non-additive joint effects

2. Concentrated Query-Time Read Pathways

InterventionResult
Donor-state swapStrong shift toward donor answer
Donor-output patchClosely reproduces state-swap effect
Output restore (after donor-state swap)Donor effect largely disappears
Control interventionsNear zero

This confirms a linked write–retain–read mechanism: content enters state at fact time, persists through context, and is expressed via concentrated query-time read pathways.

3. Selective Multi-Fact Recall

In four-fact contexts:

  • Donor answers transferred in 61–89% of cases
  • Non-target answers broken in only 0.4–1.4%
  • Cross-fact decomposition: projecting fact A's state difference onto fact B's direction produces opposing effects on B, while orthogonal components and random controls have smaller effects—demonstrating selective access coexists with cross-fact causal coupling

4. Hybrid Architecture Shift

ModelRecurrent/Local RecoveryFull-Attention KV Recovery
Pure GDN-340M~100%N/A
Hybrid GDN-340M<1%>99%
Qwen3.5-4B (3 conditions)0.2–3.6%96–100%
Qwen3.5-9B retrieval~0%~100%

This holds across additional tasks: entity updates with WikiText/PG19 filler, and RULER single-NIAH at ~4K tokens (KV recovery: 96.2–97.2%).

5. Memory Load vs. Elapsed Context

  • Increasing active associations from K=1K=1 to K=16K=16 reduces recall by 30–40 percentage points at fixed sequence length
  • Extending fact-to-query distance from 128 to 1024 tokens at fixed K=4K=4 causes only modest decline
  • Alignment of later writes: Aligning a distractor's write key with the target direction lowers target margin by 0.27–0.44 logits; orthogonalization has little effect

6. Distinct Failure Modes Under Load

MetricTrend with increasing K
Clean recallDeclines (47–60% at K=32 for 340M models; 69% for GDN-1.3B)
Target edit strengthFalls by roughly half or more
Collateral damageGDN models <2% break rate through K=32; GLA-340M exceeds 5% criterion from low load

The delta rule in GDN may suppress accumulated cross-talk and preserve selectivity better than GLA's pure gating.

Theoretical and Practical Implications

Theoretical implications:

  • Recurrent state should be viewed as a structured shared memory rather than a compressed collection of token-specific KV entries
  • The write–retain–read lifecycle provides a causal account of how fixed-size memory organizes, accesses, and loses past information
  • The architecture-dependent nature of memory substrate (recurrent vs. KV) challenges assumptions about where information "lives" in hybrid models

Practical implications for efficient long-context systems:

  • Fixed-size recurrent state offers an attractive alternative to growing KV caches for streaming assistants, long-running agents, and resource-constrained inference
  • Effectiveness depends on controlling competition among accumulated memories
  • Motivates mechanisms for selectively gating, refreshing, or partitioning recurrent writes
  • Hybrid systems could dynamically choose which information remains in recurrent state vs. token-addressable KV memory

Key limitation caveats:

  • Focus on controlled associative recall may not cover all forms of long-context reasoning
  • Pure–hybrid comparisons rely on separately trained released models
  • Concentrated pathways interact with surrounding recurrent computation and should not be interpreted as isolated memory locations

Conclusion

The paper provides a causal account of how fixed-size recurrent memory operates in pretrained linear-attention models:

  1. Remembered information enters through content-specific writes at fact time
  2. It is later accessed through concentrated query-time read pathways
  3. Multiple facts can remain selectively usable despite sharing a common state
  4. This shared organization provides fixed-size memory efficiency but introduces interference whose severity depends on memory load and overlapping writes
  5. In hybrid architectures, the direct runtime memory substrate shifts predominantly to full-attention KV memory

Future directions:

  • Translating mechanistic findings into memory management strategies for practical systems
  • Mechanisms that selectively gate, refresh, or partition recurrent writes
  • Hybrid systems that dynamically route information between recurrent state and KV memory
  • Extending analysis to application-driven settings to inform design principles for efficient long-context models

The findings clarify both the capabilities and limitations of recurrent state as a memory substrate, distinguishing it fundamentally from token-addressable KV memory while revealing the interference and capacity constraints that emerge from shared storage.

Related papers