# How Linear Attention Remembers

> Linear attention memory stores facts via concentrated content-specific writes and retrieves them via focused query-time reads, with hybrid models shifting recall to full-attention KV caches.

- **Source:** [arXiv](https://arxiv.org/abs/2609.33093)
- **Published:** 2026-10-03
- **Permalink:** https://picx.dev/p/K11iqj
- **Whiteboard:** https://picx.dev/p/K11iqj/image

## Summary

## Summary (Overview)

- **Core Question**: How does fixed-size recurrent state in linear attention architectures function as a memory system for storing, retaining, and retrieving factual information?
- **Key Finding**: Linear attention memory operates through concentrated, content-specific write pathways at fact time and concentrated query-time read pathways at recall, rather than diffuse or token-addressable storage.
- **Selective Access with Coupling**: Multiple facts can remain selectively accessible within a shared recurrent state, but exhibit cross-fact causal coupling rather than independent KV-like storage.
- **Capacity Limits**: Recall and targeted editability degrade with increasing memory load, while elapsed context alone has substantially smaller effects; interference is shaped by overlap between subsequent writes.
- **Hybrid Architecture Shift**: In hybrid models combining recurrent layers with full attention, the runtime memory substrate directly supporting recall shifts predominantly to the full-attention KV cache.

## Introduction and Theoretical Foundation

The paper addresses a fundamental tension in efficient language model design. Standard attention retains key-value (KV) representations for every preceding token, causing memory to grow linearly with sequence length—a critical bottleneck for long-context inference in both large-scale serving and resource-constrained settings like mobile devices and robotics.

Linear attention offers an alternative: instead of retaining per-token KV representations, it summarizes past information in a **fixed-size recurrent state** whose size does not grow with sequence length. However, this efficiency fundamentally changes how information is represented:

> "In a conventional KV cache, representations associated with individual tokens remain explicitly available to subsequent attention operations. In recurrent linear attention, information from many past tokens is continually updated and accumulated within the same finite state."

The paper poses several key questions:
- Which state updates carry remembered content?
- How is that information accessed at recall time?
- How do multiple facts coexist within a shared state?
- What causes retention to degrade as more information accumulates?
- Does recurrent state remain the primary memory substrate when token-addressable full attention is available?

The theoretical foundation builds on prior work connecting linear attention to **fast-weight programmers** and **associative memory interpretations** (Schlag et al., 2021; Irie et al., 2021), with architectural improvements through data-dependent gating (GLA), delta-rule updates (DeltaNet), and their combination (Gated DeltaNet/GDN).

## Methodology

### Analytical Framework

The paper develops an analytical decomposition of recurrent state memory. For a recurrent attention head, the state update is:

$$S_t = T_t(S_{t-1}) + W_t, \quad S_t \in \mathbb{R}^{d_k \times d_v}$$

where $T_t$ is linear in the previous state (e.g., $T_t(S) = G_t \odot S$ for multiplicative gating), and a common rank-one write takes the form $W_t = u_t v_t^\top$.

The read operation is:

$$o_t = q_t^\top S_t, \quad q_t \in \mathbb{R}^{d_k}, \quad o_t \in \mathbb{R}^{d_v}$$

The state can be decomposed into contributions from earlier positions:

$$S_t = \sum_{i=1}^{t} C_i^{(t)}$$

where $C_i^{(t)} = (T_t \circ T_{t-1} \circ \cdots \circ T_{i+1})(C_i^{(i)})$ and $C_i^{(i)} = W_i$.

Key theoretical insights from this decomposition:
1. **Past information shares a common state**—contributions from different positions coexist without explicit token indexing
2. **Storage and access are distinct**—a contribution can remain in $S_t$ without affecting current output
3. **The aggregate state does not uniquely reveal source contributions**—the mapping from individual contributions to their sum is many-to-one

### Experimental Design

**Models Evaluated:**
- GLA-340M (Gated Linear Attention, 341.7M params)
- GDN-340M (Gated DeltaNet, 399.5M params)
- GDN-1.3B (1.466B params)
- Hybrid GDN-340M (near-matched, 8 of 24 blocks replaced with full attention)
- Qwen3.5-4B and Qwen3.5-9B (larger hybrids)

**Task**: Controlled associative recall using WikiText-103 passages with key–value associations embedded in natural text, evaluated via four-choice recall accuracy and target logit margin.

**Causal Intervention Framework** (summarized in Figure 1 of the paper):
1. **Fact-time interventions**: Block or replace the recipient's fact-time write with donor counterpart
2. **State evolution**: Allow state to evolve under subsequent tokens
3. **Pre-query state interventions**: Swap recipient's pre-query state with donor state
4. **Query-time read interventions**: Patch or restore the selected head's output
5. **Multi-fact extension**: Test selective recall with multiple facts sharing recurrent state

**Head Selection**: Development-set screening of all recurrent heads (96/96/192 across models) using donor-state swaps, measured by the score:

$$J_{\ell,h} = \frac{1}{16}\sum_{i=1}^{16}\left[(z^{\ell,h}_{i,d_i} - z^{\ell,h}_{i,r_i}) \cdot (z^{\text{clean}}_{i,d_i} - z^{\text{clean}}_{i,r_i})\right]_i$$

Selected heads: L21H3 (GLA-340M), L20H2 (GDN-340M), L20H3 (GDN-1.3B).

## Empirical Validation / Results

### 1. Content-Specific Writes Persist in Recurrent State

- Blocking the fact-time write reduces recall by **36–79 percentage points** across models
- Replacing the recipient's update with donor counterpart makes the donor value the answer in **68–90% of cases**
- Control interventions (norm-matched random, control-head updates) do not reproduce this effect
- The selected head is part of an **interacting recurrent circuit** rather than an isolated fact store—blocking the hub head alone can impair recall more than blocking all recurrent heads together, with strongly non-additive joint effects

### 2. Concentrated Query-Time Read Pathways

| Intervention | Result |
|---|---|
| Donor-state swap | Strong shift toward donor answer |
| Donor-output patch | Closely reproduces state-swap effect |
| Output restore (after donor-state swap) | Donor effect largely disappears |
| Control interventions | Near zero |

This confirms a linked **write–retain–read mechanism**: content enters state at fact time, persists through context, and is expressed via concentrated query-time read pathways.

### 3. Selective Multi-Fact Recall

In four-fact contexts:
- Donor answers transferred in **61–89% of cases**
- Non-target answers broken in only **0.4–1.4%**
- Cross-fact decomposition: projecting fact A's state difference onto fact B's direction produces **opposing effects** on B, while orthogonal components and random controls have smaller effects—demonstrating selective access coexists with cross-fact causal coupling

### 4. Hybrid Architecture Shift

| Model | Recurrent/Local Recovery | Full-Attention KV Recovery |
|---|---|---|
| Pure GDN-340M | ~100% | N/A |
| Hybrid GDN-340M | <1% | >99% |
| Qwen3.5-4B (3 conditions) | 0.2–3.6% | 96–100% |
| Qwen3.5-9B retrieval | ~0% | ~100% |

This holds across additional tasks: entity updates with WikiText/PG19 filler, and RULER single-NIAH at ~4K tokens (KV recovery: 96.2–97.2%).

### 5. Memory Load vs. Elapsed Context

- Increasing active associations from $K=1$ to $K=16$ reduces recall by **30–40 percentage points** at fixed sequence length
- Extending fact-to-query distance from 128 to 1024 tokens at fixed $K=4$ causes only modest decline
- **Alignment of later writes**: Aligning a distractor's write key with the target direction lowers target margin by **0.27–0.44 logits**; orthogonalization has little effect

### 6. Distinct Failure Modes Under Load

| Metric | Trend with increasing K |
|---|---|
| Clean recall | Declines (47–60% at K=32 for 340M models; 69% for GDN-1.3B) |
| Target edit strength | Falls by roughly half or more |
| Collateral damage | GDN models <2% break rate through K=32; GLA-340M exceeds 5% criterion from low load |

The delta rule in GDN may suppress accumulated cross-talk and preserve selectivity better than GLA's pure gating.

## Theoretical and Practical Implications

**Theoretical implications:**
- Recurrent state should be viewed as a **structured shared memory** rather than a compressed collection of token-specific KV entries
- The write–retain–read lifecycle provides a causal account of how fixed-size memory organizes, accesses, and loses past information
- The architecture-dependent nature of memory substrate (recurrent vs. KV) challenges assumptions about where information "lives" in hybrid models

**Practical implications for efficient long-context systems:**
- Fixed-size recurrent state offers an attractive alternative to growing KV caches for streaming assistants, long-running agents, and resource-constrained inference
- Effectiveness depends on **controlling competition** among accumulated memories
- Motivates mechanisms for selectively gating, refreshing, or partitioning recurrent writes
- Hybrid systems could dynamically choose which information remains in recurrent state vs. token-addressable KV memory

**Key limitation caveats:**
- Focus on controlled associative recall may not cover all forms of long-context reasoning
- Pure–hybrid comparisons rely on separately trained released models
- Concentrated pathways interact with surrounding recurrent computation and should not be interpreted as isolated memory locations

## Conclusion

The paper provides a causal account of how fixed-size recurrent memory operates in pretrained linear-attention models:

1. Remembered information enters through **content-specific writes** at fact time
2. It is later accessed through **concentrated query-time read pathways**
3. Multiple facts can remain selectively usable despite sharing a common state
4. This shared organization provides fixed-size memory efficiency but introduces **interference** whose severity depends on memory load and overlapping writes
5. In hybrid architectures, the direct runtime memory substrate shifts predominantly to full-attention KV memory

**Future directions:**
- Translating mechanistic findings into memory management strategies for practical systems
- Mechanisms that selectively gate, refresh, or partition recurrent writes
- Hybrid systems that dynamically route information between recurrent state and KV memory
- Extending analysis to application-driven settings to inform design principles for efficient long-context models

The findings clarify both the capabilities and limitations of recurrent state as a memory substrate, distinguishing it fundamentally from token-addressable KV memory while revealing the interference and capacity constraints that emerge from shared storage.

---

_Markdown view of https://picx.dev/p/K11iqj, served by PicX — AI-generated visual whiteboard summaries of research papers._
