Full text not available for this paper

Anatomy of Associative Recall in Fixed-State Recurrences: A Matched-State Decomposition, an Interference Wall, and a Curriculum That Breaks It

Summary (Overview)

  • Controlled decomposition of recall: The paper decomposes masked multi-query associative recall (MQAR) at a fixed state budget (1,024 state elements) along three single-knob axes: short causal convolution, state-transition structure (rank-1 delta rule vs. diagonal), and decay.
  • The convolution is the dominant lever: Adding a short causal convolution improves recall by roughly +0.45–0.47 accuracy in both delta-rule and diagonal cell families, transferring across architectures and corroborating prior intervention evidence that Mamba's recall is convolution-borne.
  • The rank-1 transition advantage shrinks with convolution: The rank-1 transition beats its diagonal ablation by +0.19/+0.32 at 16/32 pairs without convolution, but the margin collapses to +0.034 once both cells carry the convolution—and a state-matched Mamba-2 ties the rank-1 cell, so no architecture-class claim survives.
  • An "interference wall" is a training-coverage gap, not a capacity limit: Cells that solve 32-pair recall fail at chance on 4-pair retrieval with distractors, flat across sequence lengths and transition types. A distance curriculum lifts the unchanged architecture from 0.021 to 1.000, revealing a "lock-in lottery" where training is stochastic and curriculum shape controls the rate.
  • Arming for recall is free on state tracking: The convolution-equipped cell is significantly better on an S₅ state-tracking guardrail at every probed depth (p ≤ 0.0044), showing no trade-off between recall capability and state-tracking learnability.

Introduction and Theoretical Foundation

Background and Motivation

Associative recall—binding key–value pairs in context and retrieving a value when its key reappears—is the capability axis on which fixed-state sequence mixers (linear attention and state-space models) most visibly trail attention. Prior work established that synthetic multi-query associative recall (MQAR) separates attention from efficient mixers and predicts much of the language-modeling gap between them (Arora et al., 2023; 2024a).

However, a modern recurrent cell is not one mechanism. A Mamba-2 block combines:

  • A diagonal selective state transition
  • An input-dependent gate
  • A short causal depthwise convolution

A Gated DeltaNet block combines:

  • A rank-1 delta-rule transition
  • A learned decay

Whole-architecture comparisons confound at least three design axes, making summary judgments ("DeltaNet-style cells recall well," "Mamba recalls better than RNNs") averages over independently toggleable knobs.

Theoretical Foundation

The paper builds on several theoretical threads:

  1. MQAR as a diagnostic task: Synthetic multi-query associative recall separates attention from efficient mixers (Arora et al., 2023; 2024a) and remains the standard lens on what a bounded recurrent state can retain (Jelassi et al., 2024; Hsieh et al., 2024).

  2. Causal intervention evidence: Mamba's induction behavior appears to live in its short convolution rather than its state-space scan (Arora et al., 2025; Parnichkun et al., 2025), though this attribution is contested once learning rates are tuned (Okpekpe & Orvieto, 2025).

  3. The delta rule: The delta rule (Schlag et al., 2021; Yang et al., 2024b) and its gated variant (Yang et al., 2024a) explicitly target recall via key-conditioned replacement; RWKV-7 generalizes the transition to diagonal-plus-rank-1 (Peng et al., 2025).

  4. Just-Read-Twice (JRT): Arora et al. (2024b) showed causal recurrent LMs recover most of the recall gap when the query precedes the context—either by repeating the prompt (JRT-Prompt) or with a bespoke prefix-LM architecture (JRT-RNN).

Methodology

The Cell Zoo

All cells are implemented in one harness as sequence mixers with identical embedding, readout, and training loops. Per head, a matrix state St∈Rdv×dkS_t \in \mathbb{R}^{d_v \times d_k} is updated by a cell-specific transition.

DeltaNet (rank-1, no decay) applies the delta rule with input-dependent write strength βt∈(0,1)\beta_t \in (0,1):

St=St−1(I−βtktkt⊤)+βtvtkt⊤,ot=Stqt(1)S_t = S_{t-1}\left(I - \beta_t k_t k_t^\top\right) + \beta_t v_t k_t^\top, \quad o_t = S_t q_t \tag{1}

Gated DeltaNet (rank-1 + decay) adds a learned scalar gate αt∈(0,1)\alpha_t \in (0,1):

St=αtSt−1(I−βtktkt⊤)+βtvtkt⊤(2)S_t = \alpha_t S_{t-1}\left(I - \beta_t k_t k_t^\top\right) + \beta_t v_t k_t^\top \tag{2}

RWKV-7 reference (rank-1 + vector decay) uses the diagonal-plus-rank-1 transition:

At=diag(wt)−(at⊙κ^t)κ^t⊤(3)A_t = \text{diag}(w_t) - (a_t \odot \hat{\kappa}_t)\hat{\kappa}_t^\top \tag{3} St=St−1At+vtkt⊤,ot=Strt(4)S_t = S_{t-1} A_t + v_t k_t^\top, \quad o_t = S_t r_t \tag{4}

with data-dependent vector decay wt∈(0,1)dkw_t \in (0,1)^{d_k}, removal key κ^t\hat{\kappa}_t, and in-context rate ata_t. Its diagonal ablation sets the rank-1 term to zero (at≡0a_t \equiv 0).

Mamba-2 reference is a pure-PyTorch state-matched implementation of the Mamba-2 SSD recurrence:

Ht=atHt−1+xtbt⊤,yt=Htct(5)H_t = a_t H_{t-1} + x_t b_t^\top, \quad y_t = H_t c_t \tag{5}

with input-dependent (at,bt,ct)(a_t, b_t, c_t) and the standard short causal depthwise convolution.

The arming knob: A width-4, strictly causal (left-padded), depthwise convolution on the interaction projections, applied pre-activation exactly as in Mamba-2. The convolution adds ≈0.5k parameters and zero recurrent state, so state-matching is preserved.

State Matching

The factorial runs at constrained state (d=32d=32, one layer). The rank-1 cells carry dk2=322=1024d_k^2 = 32^2 = 1024 recurrent state elements per head group; the Mamba-2 reference is set to dstate=16d_{\text{state}} = 16 so that dinner⋅dstate=64⋅16=1024d_{\text{inner}} \cdot d_{\text{state}} = 64 \cdot 16 = 1024 elements—matched state and matched parameters (≈28.9k vs. ≈29.3k).

Tasks

  1. Masked MQAR: Each sequence is a table of KK key–value pairs followed by queried keys whose answers are single masked tokens. Values are drawn fresh per sequence from a 26-token range. K∈{8,16,32}K \in \{8, 16, 32\}. Floors: value-marginal ≈ 0.038; strongest no-binding strategy ≈ 0.068.

  2. Haystack retrieval: 4 key–value pairs, then a distractor haystack, then the query; answer is a single masked token at the final position. Lengths L∈{64,…,512}L \in \{64, \ldots, 512\}; chance ≈ 0.019. The query-at-end layout makes the task 2-hop; all arms get 2 layers.

  3. Collision-key variant: Decoy pairs in the gap reuse the table's own keys with fresh random values; ground truth is the first (table) binding. This defeats any lexical write gate.

  4. S₅ guardrail: A running-product word problem over the symmetric group S5S_5; generators drawn uniformly from the full group (answer marginal = 1/120 ≈ 0.008 at every depth); all product slots masked simultaneously.

Statistical Protocol

  • Exploratory sweeps: 3 seeds
  • Headline ablations: 10 seeds, two-sided permutation tests on mean gaps
  • Rebinding control: Evaluates the trained model on sequences whose key–value bindings have been deranged, scoring the original value. A genuine retriever's rebinding score must collapse toward the floors.
  • Lock-in rates: Where seed outcomes are bimodal (a seed either "locks in" or does not), report fraction of seeds above threshold with Fisher exact tests rather than means.

Empirical Validation / Results

The Matched-State Decomposition

Table 1: Constrained state, K=32 (hardest setting)

ArmRecall mean [min, max]Ingredients
armed gdn0.983 [0.917, 1.000]rank-1 + decay + conv
dg rank (RWKV-7 rank-1)0.653 [0.507, 0.874]rank-1 (vector decay), no conv
mamba2 ref (state-matched)0.592 [0.196, 0.884]diagonal + selectivity + conv
deltanet0.561 [0.094, 0.927]rank-1, no decay, no conv
gated deltanet0.518 [0.296, 0.798]rank-1 + decay, no conv
dg diag (RWKV-7, rank off)0.390 [0.251, 0.691]diagonal (vector decay)
mamba2 ref noconv0.151 [0.079, 0.191]diagonal + selectivity, no conv

Axis 1: The convolution (≈ +0.45, family-transferable)

  • Delta-rule cell: 0.518 → 0.983 (+0.47; gated deltanet → armed gdn)
  • Diagonal cell: 0.151 → 0.592 (+0.44; mamba2 ref noconv → mamba2 ref)
  • The convolution is a state-free local primitive, yet the largest single lever in the factorial.

Axis 2: The transition (≈ +0.3 without convolution; ≈ +0.03 with it)

  • Without convolution: dg rank vs. dg diag = 0.653 vs. 0.390 (+0.26); hardened: +0.191 at K=16 (p=0.0002), +0.319 at K=32 (p<10⁻⁴)
  • With convolution: 0.992 (rank-1) vs. 0.958 (diagonal), +0.034 (p=0.0023) at K=32

Axis 3: Decay (no measurable cost)

  • deltanet (no decay) 0.561 vs. gated deltanet (decay) 0.518 at K=32: p=0.86 at n=20
  • Decay appears to trade occasional lucky solves for consistency; neither is a mean-level effect.

The Hardened Transition Ablation

Table 3: Hardened transition ablation (n=10 seeds)

KArmMean [95% CI]Lock-in ≥ 0.9ReboundDrop
16dg rank (rank-1)0.998 [0.995, 0.999]10/100.024+0.974
16dg diag (diagonal ablation)0.806 [0.721, 0.892]3/100.038+0.769
16deltanet (rank-1)1.000 [0.999, 1.000]10/100.024+0.976
16mamba2 ds16 (diagonal, matched)0.999 [0.999, 1.000]10/100.024+0.975
32mamba2 ds16 (diagonal, matched)0.849 [0.807, 0.890]3/100.042+0.807
32dg rank (rank-1)0.702 [0.631, 0.769]0/100.051+0.651
32deltanet (rank-1)0.623 [0.482, 0.741]1/100.057+0.566
32dg diag (diagonal ablation)0.382 [0.322, 0.450]0/100.073+0.309

Key findings:

  1. The within-cell ablation replicates decisively: dg rank > dg diag
  2. No class claim survives: A state- and parameter-matched Mamba-2 ties the rank-1 cell at K=16 (p=0.25) and beats it at K=32 (0.849 vs. 0.702, p=0.0036)
  3. The two statements together resolve the folklore: convolution-free recurrent cells lose to Mamba because of the convolution, not the recurrence

The Interference Wall

Table 4: The haystack wall, flat in length (n=5 seeds per cell)

Arm (2 layers)L=64L=128L=256
attention (head dim 16)1.0001.0001.000
armed gdn0.0180.0180.019
mamba2 ref0.0170.0190.019
dg rank (RWKV-7)0.0180.0180.018

Four discriminating properties:

  1. Not capacity: Same cells store 32 competing bindings at ≥0.99 but fail on 4 bindings plus distractors
  2. A wall, not a dip: Failure is total at the shortest tested length, inside the training distribution, flat through 4× the length
  3. Transition richness does not help: GDN's delta rule, Mamba's selectivity, and RWKV-7's removal key wall identically
  4. Attention is untouched: Per-token KV entries cannot be overwritten by later writes

The Curriculum Fix

Table 5: Curriculum lever attribution (n=10 seeds, L=64 haystack, 2 layers)

CellDenseCurric.Lock-inPer-seed
armed gdn (bidir)4✓6/101.00/0.87/0.88/0.99/0.36/0.45/0.12/0.04/1.00/0.97
armed gdn (bidir)1✓7/100.82/0.13/0.06/1.00/0.92/1.00/0.58/1.00/1.00/1.00
armed gdn (bidir)4—1/100.02/0.02/0.02/1.00/0.02/0.02/0.02/0.02/0.02/0.02
armed gdn fwd (causal)4✓5/100.84/0.75/0.96/1.00/0.06/0.83/0.62/0.24/0.08/0.96
attention (abs. pos.)4✓0/5all ≈ 0.046
attention (rotary)4✓6/10mean 0.84; range 0.28–1.00

Key results:

  • The distance curriculum lifts the walled cell from 0.021 (chance) to 1.000—an existence proof by construction
  • Lock-in rises from 1/10 seeds (dense only) to 7/10 (curriculum), p=0.02 (Fisher two-sided)
  • Adding dense supervision to the curriculum changes nothing (6/10)
  • Length boundary: At L=256, uniform curriculum fails everywhere (0/9 trials) but a shaped ramp locks in 4/5 seeds (p≈0.005); at L=512, the time-based ramp collapses (0/6) while a success-gated ramp locks in 6/6 (p=0.0011)

Bidirectional Denoisers

Table 6: Collision-key discriminator (L=64, dense-4 + curriculum, 2 layers, n=10 seeds)

ArmLock-inMeanPer-seed range
armed gdn (bidirectional)3/100.3830.023–1.000
armed gdn fwd (causal)1/100.3010.036–0.844
gated deltanet (bidir, no conv)0/100.1280.017–0.238

Key findings:

  • No directional margin at ten seeds (mean difference +0.082, p=0.63)
  • Both directions can solve the task; any solving circuit needs two layers (mark-and-route)
  • The convolution-free cell never locks in even under the shaped curriculum (0/10 vs. 9/10, p≈10⁻⁴)
  • Curriculum shape is a powerful direction-agnostic stabilizer (1/10 → 9/10, p≈6×10⁻⁴)

Guardrail: Arming Is Free on State Tracking

Table 7: S₅ guardrail (d=128; chance ≈ 0.008)

S₅ deptharmed gdngated deltanetp
L=2 (n=10)1.000 [1.000, 1.000]0.980 [0.945, 1.000]0.0002
L=4 (n=10)0.524 [0.492, 0.637]0.389 [0.274, 0.514]0.0014
L=8 (n=10)0.214 [0.143, 0.261]0.152 [0.132, 0.240]0.0044
L=16 (n=10)0.087 [0.072, 0.128]0.071 [0.070, 0.073]<10⁻⁴

The armed cell is significantly better at every probed depth: arming for recall does not trade away state tracking.

Theoretical and Practical Implications

Re-scoping "Recurrent Models Are Bad at Recall"

At matched state, the claim decomposes into five things:

  1. A missing convolution (large, fixable for free in state terms)
  2. A transition effect (real, within-cell)
  3. A decay tax (real, scalar-gate-specific)
  4. An interference failure under sparse supervision (fixable by curriculum, within a length boundary)
  5. A directionality property (bidirectional denoisers get query-first reading natively)

None of these is "the recurrence cannot bind."

When Is a Hybrid Actually Needed?

Two classic triggers for "add attention layers for retrieval" dissolve under cheap interventions:

  • The unarmed-cell recall deficit (add the convolution)
  • The within-capacity haystack wall (train with a distance curriculum)

Three things remain as genuine hybrid territory:

  • Loads beyond state capacity (the graceful K-limit)
  • Lengths beyond the curriculum boundary (L ≥ 256 here, pending shaped curricula)
  • Query-conditioned regimes a fixed-state writer cannot anticipate

For Recurrent Diffusion LMs

The JRT-for-free result gives bidirectional recurrent denoisers a principled recall story that causal recurrent LMs lack: the objective itself buys the second read. The requirements are concrete—two layers minimum and a short convolution in the backward stream.

The Central Conjecture

Within the capacity regime of a fixed-state recurrence, apparent retrieval walls are training-coverage gaps.

The length boundary at L=512 is its first direct test: the cell that sits at 0/6 under a time-based ramp locks in 6/6 once the ramp is gated on measured competence, so that boundary was the schedule, not the circuit.

Conclusion

Main Takeaways

  1. The convolution is the recall story at constrained state: An armed diagonal cell reaches 0.958, showing that at constrained state the short convolution is very nearly the whole recall story.

  2. No architecture-class claim survives matched-state comparison: The rank-1 transition's advantage over diagonal ablations is real but within-cell; a state-matched Mamba-2 ties the rank-1 cell, so "rank-1 > diagonal" as a class claim is false.

  3. The interference wall is a training-coverage gap: The same architecture that sits at chance under naive training reaches 1.000 under a distance curriculum—an existence proof that recurrent retrieval failure is substantially a learnability problem.

  4. Training is a lock-in lottery: Seed outcomes are bimodal (a seed either locks in or does not), and curriculum shape is the lever that controls the rate. Success-gated ramps reopen boundaries that time-based schedules cannot.

  5. Arming for recall is free: The convolution-equipped cell is significantly better on S₅ state tracking at every depth, showing no trade-off between recall capability and state-tracking learnability.

Future Directions

  • Whether some length defeats the gated schedule too (testing the conjecture's boundary)
  • Whether the query-first mechanism carries to scale (untested beyond toy scale)
  • Whether context-modulated forgetting (Kimi Team, 2025) becomes a useful design axis with better-powered tests
  • The companion paper (Boesch & Wee, 2026) takes the curriculum half to converted diffusion LMs at 1.7B and 8B parameters

Limitations

  • Toy scale (d∈{32,128}d \in \{32, 128\}, synthetic tasks)
  • The Mamba-2 reference under-trains the official kernel
  • Learning-rate sensitivity affects magnitude claims
  • The attention control is valid only where reported (two-layer regimes)
  • Cross-family magnitude comparisons carry a parameterization caveat

Related papers