# Anatomy of Associative Recall in Fixed-State Recurrences: A Matched-State Decomposition, an Interference Wall, and a Curriculum That Breaks It

> At matched state, recall in recurrent cells is dominated by a short convolution, while an apparent interference wall is a training-coverage gap fixable by curriculum.

- **Source:** [arXiv](https://arxiv.org/abs/2609.16183)
- **Published:** 2026-10-03
- **Permalink:** https://picx.dev/p/Jcsu3h
- **Whiteboard:** https://picx.dev/p/Jcsu3h/image

## Summary

# Anatomy of Associative Recall in Fixed-State Recurrences: A Matched-State Decomposition, an Interference Wall, and a Curriculum That Breaks It

## Summary (Overview)

- **Controlled decomposition of recall**: The paper decomposes masked multi-query associative recall (MQAR) at a fixed state budget (1,024 state elements) along three single-knob axes: short causal convolution, state-transition structure (rank-1 delta rule vs. diagonal), and decay.
- **The convolution is the dominant lever**: Adding a short causal convolution improves recall by roughly +0.45–0.47 accuracy in both delta-rule and diagonal cell families, transferring across architectures and corroborating prior intervention evidence that Mamba's recall is convolution-borne.
- **The rank-1 transition advantage shrinks with convolution**: The rank-1 transition beats its diagonal ablation by +0.19/+0.32 at 16/32 pairs without convolution, but the margin collapses to +0.034 once both cells carry the convolution—and a state-matched Mamba-2 ties the rank-1 cell, so no architecture-class claim survives.
- **An "interference wall" is a training-coverage gap, not a capacity limit**: Cells that solve 32-pair recall fail at chance on 4-pair retrieval with distractors, flat across sequence lengths and transition types. A distance curriculum lifts the unchanged architecture from 0.021 to 1.000, revealing a "lock-in lottery" where training is stochastic and curriculum shape controls the rate.
- **Arming for recall is free on state tracking**: The convolution-equipped cell is significantly better on an S₅ state-tracking guardrail at every probed depth (p ≤ 0.0044), showing no trade-off between recall capability and state-tracking learnability.

## Introduction and Theoretical Foundation

### Background and Motivation

Associative recall—binding key–value pairs in context and retrieving a value when its key reappears—is the capability axis on which fixed-state sequence mixers (linear attention and state-space models) most visibly trail attention. Prior work established that synthetic multi-query associative recall (MQAR) separates attention from efficient mixers and predicts much of the language-modeling gap between them (Arora et al., 2023; 2024a).

However, a modern recurrent cell is not one mechanism. A Mamba-2 block combines:
- A diagonal selective state transition
- An input-dependent gate
- A short causal depthwise convolution

A Gated DeltaNet block combines:
- A rank-1 delta-rule transition
- A learned decay

Whole-architecture comparisons confound at least three design axes, making summary judgments ("DeltaNet-style cells recall well," "Mamba recalls better than RNNs") averages over independently toggleable knobs.

### Theoretical Foundation

The paper builds on several theoretical threads:

1. **MQAR as a diagnostic task**: Synthetic multi-query associative recall separates attention from efficient mixers (Arora et al., 2023; 2024a) and remains the standard lens on what a bounded recurrent state can retain (Jelassi et al., 2024; Hsieh et al., 2024).

2. **Causal intervention evidence**: Mamba's induction behavior appears to live in its short convolution rather than its state-space scan (Arora et al., 2025; Parnichkun et al., 2025), though this attribution is contested once learning rates are tuned (Okpekpe & Orvieto, 2025).

3. **The delta rule**: The delta rule (Schlag et al., 2021; Yang et al., 2024b) and its gated variant (Yang et al., 2024a) explicitly target recall via key-conditioned replacement; RWKV-7 generalizes the transition to diagonal-plus-rank-1 (Peng et al., 2025).

4. **Just-Read-Twice (JRT)**: Arora et al. (2024b) showed causal recurrent LMs recover most of the recall gap when the query precedes the context—either by repeating the prompt (JRT-Prompt) or with a bespoke prefix-LM architecture (JRT-RNN).

## Methodology

### The Cell Zoo

All cells are implemented in one harness as sequence mixers with identical embedding, readout, and training loops. Per head, a matrix state $S_t \in \mathbb{R}^{d_v \times d_k}$ is updated by a cell-specific transition.

**DeltaNet (rank-1, no decay)** applies the delta rule with input-dependent write strength $\beta_t \in (0,1)$:

$$S_t = S_{t-1}\left(I - \beta_t k_t k_t^\top\right) + \beta_t v_t k_t^\top, \quad o_t = S_t q_t \tag{1}$$

**Gated DeltaNet (rank-1 + decay)** adds a learned scalar gate $\alpha_t \in (0,1)$:

$$S_t = \alpha_t S_{t-1}\left(I - \beta_t k_t k_t^\top\right) + \beta_t v_t k_t^\top \tag{2}$$

**RWKV-7 reference (rank-1 + vector decay)** uses the diagonal-plus-rank-1 transition:

$$A_t = \text{diag}(w_t) - (a_t \odot \hat{\kappa}_t)\hat{\kappa}_t^\top \tag{3}$$

$$S_t = S_{t-1} A_t + v_t k_t^\top, \quad o_t = S_t r_t \tag{4}$$

with data-dependent vector decay $w_t \in (0,1)^{d_k}$, removal key $\hat{\kappa}_t$, and in-context rate $a_t$. Its diagonal ablation sets the rank-1 term to zero ($a_t \equiv 0$).

**Mamba-2 reference** is a pure-PyTorch state-matched implementation of the Mamba-2 SSD recurrence:

$$H_t = a_t H_{t-1} + x_t b_t^\top, \quad y_t = H_t c_t \tag{5}$$

with input-dependent $(a_t, b_t, c_t)$ and the standard short causal depthwise convolution.

**The arming knob**: A width-4, strictly causal (left-padded), depthwise convolution on the interaction projections, applied pre-activation exactly as in Mamba-2. The convolution adds ≈0.5k parameters and **zero recurrent state**, so state-matching is preserved.

### State Matching

The factorial runs at constrained state ($d=32$, one layer). The rank-1 cells carry $d_k^2 = 32^2 = 1024$ recurrent state elements per head group; the Mamba-2 reference is set to $d_{\text{state}} = 16$ so that $d_{\text{inner}} \cdot d_{\text{state}} = 64 \cdot 16 = 1024$ elements—matched state and matched parameters (≈28.9k vs. ≈29.3k).

### Tasks

1. **Masked MQAR**: Each sequence is a table of $K$ key–value pairs followed by queried keys whose answers are single masked tokens. Values are drawn fresh per sequence from a 26-token range. $K \in \{8, 16, 32\}$. Floors: value-marginal ≈ 0.038; strongest no-binding strategy ≈ 0.068.

2. **Haystack retrieval**: 4 key–value pairs, then a distractor haystack, then the query; answer is a single masked token at the final position. Lengths $L \in \{64, \ldots, 512\}$; chance ≈ 0.019. The query-at-end layout makes the task 2-hop; all arms get 2 layers.

3. **Collision-key variant**: Decoy pairs in the gap reuse the table's own keys with fresh random values; ground truth is the first (table) binding. This defeats any lexical write gate.

4. **S₅ guardrail**: A running-product word problem over the symmetric group $S_5$; generators drawn uniformly from the full group (answer marginal = 1/120 ≈ 0.008 at every depth); all product slots masked simultaneously.

### Statistical Protocol

- Exploratory sweeps: 3 seeds
- Headline ablations: 10 seeds, two-sided permutation tests on mean gaps
- **Rebinding control**: Evaluates the trained model on sequences whose key–value bindings have been deranged, scoring the original value. A genuine retriever's rebinding score must collapse toward the floors.
- **Lock-in rates**: Where seed outcomes are bimodal (a seed either "locks in" or does not), report fraction of seeds above threshold with Fisher exact tests rather than means.

## Empirical Validation / Results

### The Matched-State Decomposition

**Table 1: Constrained state, K=32 (hardest setting)**

| Arm | Recall mean [min, max] | Ingredients |
|-----|------------------------|-------------|
| armed gdn | 0.983 [0.917, 1.000] | rank-1 + decay + conv |
| dg rank (RWKV-7 rank-1) | 0.653 [0.507, 0.874] | rank-1 (vector decay), no conv |
| mamba2 ref (state-matched) | 0.592 [0.196, 0.884] | diagonal + selectivity + conv |
| deltanet | 0.561 [0.094, 0.927] | rank-1, no decay, no conv |
| gated deltanet | 0.518 [0.296, 0.798] | rank-1 + decay, no conv |
| dg diag (RWKV-7, rank off) | 0.390 [0.251, 0.691] | diagonal (vector decay) |
| mamba2 ref noconv | 0.151 [0.079, 0.191] | diagonal + selectivity, no conv |

**Axis 1: The convolution (≈ +0.45, family-transferable)**
- Delta-rule cell: 0.518 → 0.983 (+0.47; gated deltanet → armed gdn)
- Diagonal cell: 0.151 → 0.592 (+0.44; mamba2 ref noconv → mamba2 ref)
- The convolution is a state-free local primitive, yet the largest single lever in the factorial.

**Axis 2: The transition (≈ +0.3 without convolution; ≈ +0.03 with it)**
- Without convolution: dg rank vs. dg diag = 0.653 vs. 0.390 (+0.26); hardened: +0.191 at K=16 (p=0.0002), +0.319 at K=32 (p<10⁻⁴)
- With convolution: 0.992 (rank-1) vs. 0.958 (diagonal), +0.034 (p=0.0023) at K=32

**Axis 3: Decay (no measurable cost)**
- deltanet (no decay) 0.561 vs. gated deltanet (decay) 0.518 at K=32: p=0.86 at n=20
- Decay appears to trade occasional lucky solves for consistency; neither is a mean-level effect.

### The Hardened Transition Ablation

**Table 3: Hardened transition ablation (n=10 seeds)**

| K | Arm | Mean [95% CI] | Lock-in ≥ 0.9 | Rebound | Drop |
|---|-----|---------------|----------------|---------|------|
| 16 | dg rank (rank-1) | 0.998 [0.995, 0.999] | 10/10 | 0.024 | +0.974 |
| 16 | dg diag (diagonal ablation) | 0.806 [0.721, 0.892] | 3/10 | 0.038 | +0.769 |
| 16 | deltanet (rank-1) | 1.000 [0.999, 1.000] | 10/10 | 0.024 | +0.976 |
| 16 | mamba2 ds16 (diagonal, matched) | 0.999 [0.999, 1.000] | 10/10 | 0.024 | +0.975 |
| 32 | mamba2 ds16 (diagonal, matched) | 0.849 [0.807, 0.890] | 3/10 | 0.042 | +0.807 |
| 32 | dg rank (rank-1) | 0.702 [0.631, 0.769] | 0/10 | 0.051 | +0.651 |
| 32 | deltanet (rank-1) | 0.623 [0.482, 0.741] | 1/10 | 0.057 | +0.566 |
| 32 | dg diag (diagonal ablation) | 0.382 [0.322, 0.450] | 0/10 | 0.073 | +0.309 |

**Key findings:**
1. The within-cell ablation replicates decisively: dg rank > dg diag
2. **No class claim survives**: A state- and parameter-matched Mamba-2 ties the rank-1 cell at K=16 (p=0.25) and beats it at K=32 (0.849 vs. 0.702, p=0.0036)
3. The two statements together resolve the folklore: convolution-free recurrent cells lose to Mamba because of the convolution, not the recurrence

### The Interference Wall

**Table 4: The haystack wall, flat in length (n=5 seeds per cell)**

| Arm (2 layers) | L=64 | L=128 | L=256 |
|----------------|------|-------|-------|
| attention (head dim 16) | 1.000 | 1.000 | 1.000 |
| armed gdn | 0.018 | 0.018 | 0.019 |
| mamba2 ref | 0.017 | 0.019 | 0.019 |
| dg rank (RWKV-7) | 0.018 | 0.018 | 0.018 |

**Four discriminating properties:**
1. **Not capacity**: Same cells store 32 competing bindings at ≥0.99 but fail on 4 bindings plus distractors
2. **A wall, not a dip**: Failure is total at the shortest tested length, inside the training distribution, flat through 4× the length
3. **Transition richness does not help**: GDN's delta rule, Mamba's selectivity, and RWKV-7's removal key wall identically
4. **Attention is untouched**: Per-token KV entries cannot be overwritten by later writes

### The Curriculum Fix

**Table 5: Curriculum lever attribution (n=10 seeds, L=64 haystack, 2 layers)**

| Cell | Dense | Curric. | Lock-in | Per-seed |
|------|-------|---------|---------|----------|
| armed gdn (bidir) | 4 | ✓ | 6/10 | 1.00/0.87/0.88/0.99/0.36/0.45/0.12/0.04/1.00/0.97 |
| armed gdn (bidir) | 1 | ✓ | 7/10 | 0.82/0.13/0.06/1.00/0.92/1.00/0.58/1.00/1.00/1.00 |
| armed gdn (bidir) | 4 | — | 1/10 | 0.02/0.02/0.02/1.00/0.02/0.02/0.02/0.02/0.02/0.02 |
| armed gdn fwd (causal) | 4 | ✓ | 5/10 | 0.84/0.75/0.96/1.00/0.06/0.83/0.62/0.24/0.08/0.96 |
| attention (abs. pos.) | 4 | ✓ | 0/5 | all ≈ 0.046 |
| attention (rotary) | 4 | ✓ | 6/10 | mean 0.84; range 0.28–1.00 |

**Key results:**
- The distance curriculum lifts the walled cell from 0.021 (chance) to 1.000—an existence proof by construction
- Lock-in rises from 1/10 seeds (dense only) to 7/10 (curriculum), p=0.02 (Fisher two-sided)
- Adding dense supervision to the curriculum changes nothing (6/10)
- **Length boundary**: At L=256, uniform curriculum fails everywhere (0/9 trials) but a shaped ramp locks in 4/5 seeds (p≈0.005); at L=512, the time-based ramp collapses (0/6) while a success-gated ramp locks in 6/6 (p=0.0011)

### Bidirectional Denoisers

**Table 6: Collision-key discriminator (L=64, dense-4 + curriculum, 2 layers, n=10 seeds)**

| Arm | Lock-in | Mean | Per-seed range |
|-----|---------|------|----------------|
| armed gdn (bidirectional) | 3/10 | 0.383 | 0.023–1.000 |
| armed gdn fwd (causal) | 1/10 | 0.301 | 0.036–0.844 |
| gated deltanet (bidir, no conv) | 0/10 | 0.128 | 0.017–0.238 |

**Key findings:**
- No directional margin at ten seeds (mean difference +0.082, p=0.63)
- Both directions can solve the task; any solving circuit needs two layers (mark-and-route)
- The convolution-free cell never locks in even under the shaped curriculum (0/10 vs. 9/10, p≈10⁻⁴)
- Curriculum shape is a powerful direction-agnostic stabilizer (1/10 → 9/10, p≈6×10⁻⁴)

### Guardrail: Arming Is Free on State Tracking

**Table 7: S₅ guardrail (d=128; chance ≈ 0.008)**

| S₅ depth | armed gdn | gated deltanet | p |
|----------|-----------|----------------|---|
| L=2 (n=10) | 1.000 [1.000, 1.000] | 0.980 [0.945, 1.000] | 0.0002 |
| L=4 (n=10) | 0.524 [0.492, 0.637] | 0.389 [0.274, 0.514] | 0.0014 |
| L=8 (n=10) | 0.214 [0.143, 0.261] | 0.152 [0.132, 0.240] | 0.0044 |
| L=16 (n=10) | 0.087 [0.072, 0.128] | 0.071 [0.070, 0.073] | <10⁻⁴ |

The armed cell is significantly better at every probed depth: arming for recall does not trade away state tracking.

## Theoretical and Practical Implications

### Re-scoping "Recurrent Models Are Bad at Recall"

At matched state, the claim decomposes into five things:
1. A **missing convolution** (large, fixable for free in state terms)
2. A **transition effect** (real, within-cell)
3. A **decay tax** (real, scalar-gate-specific)
4. An **interference failure under sparse supervision** (fixable by curriculum, within a length boundary)
5. A **directionality property** (bidirectional denoisers get query-first reading natively)

None of these is "the recurrence cannot bind."

### When Is a Hybrid Actually Needed?

Two classic triggers for "add attention layers for retrieval" dissolve under cheap interventions:
- The unarmed-cell recall deficit (add the convolution)
- The within-capacity haystack wall (train with a distance curriculum)

Three things remain as genuine hybrid territory:
- Loads beyond state capacity (the graceful K-limit)
- Lengths beyond the curriculum boundary (L ≥ 256 here, pending shaped curricula)
- Query-conditioned regimes a fixed-state writer cannot anticipate

### For Recurrent Diffusion LMs

The JRT-for-free result gives bidirectional recurrent denoisers a principled recall story that causal recurrent LMs lack: the objective itself buys the second read. The requirements are concrete—two layers minimum and a short convolution in the backward stream.

### The Central Conjecture

> Within the capacity regime of a fixed-state recurrence, apparent retrieval walls are training-coverage gaps.

The length boundary at L=512 is its first direct test: the cell that sits at 0/6 under a time-based ramp locks in 6/6 once the ramp is gated on measured competence, so that boundary was the schedule, not the circuit.

## Conclusion

### Main Takeaways

1. **The convolution is the recall story at constrained state**: An armed diagonal cell reaches 0.958, showing that at constrained state the short convolution is very nearly the whole recall story.

2. **No architecture-class claim survives matched-state comparison**: The rank-1 transition's advantage over diagonal ablations is real but within-cell; a state-matched Mamba-2 ties the rank-1 cell, so "rank-1 > diagonal" as a class claim is false.

3. **The interference wall is a training-coverage gap**: The same architecture that sits at chance under naive training reaches 1.000 under a distance curriculum—an existence proof that recurrent retrieval failure is substantially a learnability problem.

4. **Training is a lock-in lottery**: Seed outcomes are bimodal (a seed either locks in or does not), and curriculum shape is the lever that controls the rate. Success-gated ramps reopen boundaries that time-based schedules cannot.

5. **Arming for recall is free**: The convolution-equipped cell is significantly better on S₅ state tracking at every depth, showing no trade-off between recall capability and state-tracking learnability.

### Future Directions

- Whether some length defeats the gated schedule too (testing the conjecture's boundary)
- Whether the query-first mechanism carries to scale (untested beyond toy scale)
- Whether context-modulated forgetting (Kimi Team, 2025) becomes a useful design axis with better-powered tests
- The companion paper (Boesch & Wee, 2026) takes the curriculum half to converted diffusion LMs at 1.7B and 8B parameters

### Limitations

- Toy scale ($d \in \{32, 128\}$, synthetic tasks)
- The Mamba-2 reference under-trains the official kernel
- Learning-rate sensitivity affects magnitude claims
- The attention control is valid only where reported (two-layer regimes)
- Cross-family magnitude comparisons carry a parameterization caveat

---

_Markdown view of https://picx.dev/p/Jcsu3h, served by PicX — AI-generated visual whiteboard summaries of research papers._
