Full text not available for this paper
Anatomy of Associative Recall in Fixed-State Recurrences: A Matched-State Decomposition, an Interference Wall, and a Curriculum That Breaks It
Summary (Overview)
- Controlled decomposition of recall: The paper decomposes masked multi-query associative recall (MQAR) at a fixed state budget (1,024 state elements) along three single-knob axes: short causal convolution, state-transition structure (rank-1 delta rule vs. diagonal), and decay.
- The convolution is the dominant lever: Adding a short causal convolution improves recall by roughly +0.45–0.47 accuracy in both delta-rule and diagonal cell families, transferring across architectures and corroborating prior intervention evidence that Mamba's recall is convolution-borne.
- The rank-1 transition advantage shrinks with convolution: The rank-1 transition beats its diagonal ablation by +0.19/+0.32 at 16/32 pairs without convolution, but the margin collapses to +0.034 once both cells carry the convolution—and a state-matched Mamba-2 ties the rank-1 cell, so no architecture-class claim survives.
- An "interference wall" is a training-coverage gap, not a capacity limit: Cells that solve 32-pair recall fail at chance on 4-pair retrieval with distractors, flat across sequence lengths and transition types. A distance curriculum lifts the unchanged architecture from 0.021 to 1.000, revealing a "lock-in lottery" where training is stochastic and curriculum shape controls the rate.
- Arming for recall is free on state tracking: The convolution-equipped cell is significantly better on an S₅ state-tracking guardrail at every probed depth (p ≤ 0.0044), showing no trade-off between recall capability and state-tracking learnability.
Introduction and Theoretical Foundation
Background and Motivation
Associative recall—binding key–value pairs in context and retrieving a value when its key reappears—is the capability axis on which fixed-state sequence mixers (linear attention and state-space models) most visibly trail attention. Prior work established that synthetic multi-query associative recall (MQAR) separates attention from efficient mixers and predicts much of the language-modeling gap between them (Arora et al., 2023; 2024a).
However, a modern recurrent cell is not one mechanism. A Mamba-2 block combines:
- A diagonal selective state transition
- An input-dependent gate
- A short causal depthwise convolution
A Gated DeltaNet block combines:
- A rank-1 delta-rule transition
- A learned decay
Whole-architecture comparisons confound at least three design axes, making summary judgments ("DeltaNet-style cells recall well," "Mamba recalls better than RNNs") averages over independently toggleable knobs.
Theoretical Foundation
The paper builds on several theoretical threads:
-
MQAR as a diagnostic task: Synthetic multi-query associative recall separates attention from efficient mixers (Arora et al., 2023; 2024a) and remains the standard lens on what a bounded recurrent state can retain (Jelassi et al., 2024; Hsieh et al., 2024).
-
Causal intervention evidence: Mamba's induction behavior appears to live in its short convolution rather than its state-space scan (Arora et al., 2025; Parnichkun et al., 2025), though this attribution is contested once learning rates are tuned (Okpekpe & Orvieto, 2025).
-
The delta rule: The delta rule (Schlag et al., 2021; Yang et al., 2024b) and its gated variant (Yang et al., 2024a) explicitly target recall via key-conditioned replacement; RWKV-7 generalizes the transition to diagonal-plus-rank-1 (Peng et al., 2025).
-
Just-Read-Twice (JRT): Arora et al. (2024b) showed causal recurrent LMs recover most of the recall gap when the query precedes the context—either by repeating the prompt (JRT-Prompt) or with a bespoke prefix-LM architecture (JRT-RNN).
Methodology
The Cell Zoo
All cells are implemented in one harness as sequence mixers with identical embedding, readout, and training loops. Per head, a matrix state is updated by a cell-specific transition.
DeltaNet (rank-1, no decay) applies the delta rule with input-dependent write strength :
Gated DeltaNet (rank-1 + decay) adds a learned scalar gate :
RWKV-7 reference (rank-1 + vector decay) uses the diagonal-plus-rank-1 transition:
with data-dependent vector decay , removal key , and in-context rate . Its diagonal ablation sets the rank-1 term to zero ().
Mamba-2 reference is a pure-PyTorch state-matched implementation of the Mamba-2 SSD recurrence:
with input-dependent and the standard short causal depthwise convolution.
The arming knob: A width-4, strictly causal (left-padded), depthwise convolution on the interaction projections, applied pre-activation exactly as in Mamba-2. The convolution adds ≈0.5k parameters and zero recurrent state, so state-matching is preserved.
State Matching
The factorial runs at constrained state (, one layer). The rank-1 cells carry recurrent state elements per head group; the Mamba-2 reference is set to so that elements—matched state and matched parameters (≈28.9k vs. ≈29.3k).
Tasks
-
Masked MQAR: Each sequence is a table of key–value pairs followed by queried keys whose answers are single masked tokens. Values are drawn fresh per sequence from a 26-token range. . Floors: value-marginal ≈ 0.038; strongest no-binding strategy ≈ 0.068.
-
Haystack retrieval: 4 key–value pairs, then a distractor haystack, then the query; answer is a single masked token at the final position. Lengths ; chance ≈ 0.019. The query-at-end layout makes the task 2-hop; all arms get 2 layers.
-
Collision-key variant: Decoy pairs in the gap reuse the table's own keys with fresh random values; ground truth is the first (table) binding. This defeats any lexical write gate.
-
S₅ guardrail: A running-product word problem over the symmetric group ; generators drawn uniformly from the full group (answer marginal = 1/120 ≈ 0.008 at every depth); all product slots masked simultaneously.
Statistical Protocol
- Exploratory sweeps: 3 seeds
- Headline ablations: 10 seeds, two-sided permutation tests on mean gaps
- Rebinding control: Evaluates the trained model on sequences whose key–value bindings have been deranged, scoring the original value. A genuine retriever's rebinding score must collapse toward the floors.
- Lock-in rates: Where seed outcomes are bimodal (a seed either "locks in" or does not), report fraction of seeds above threshold with Fisher exact tests rather than means.
Empirical Validation / Results
The Matched-State Decomposition
Table 1: Constrained state, K=32 (hardest setting)
| Arm | Recall mean [min, max] | Ingredients |
|---|---|---|
| armed gdn | 0.983 [0.917, 1.000] | rank-1 + decay + conv |
| dg rank (RWKV-7 rank-1) | 0.653 [0.507, 0.874] | rank-1 (vector decay), no conv |
| mamba2 ref (state-matched) | 0.592 [0.196, 0.884] | diagonal + selectivity + conv |
| deltanet | 0.561 [0.094, 0.927] | rank-1, no decay, no conv |
| gated deltanet | 0.518 [0.296, 0.798] | rank-1 + decay, no conv |
| dg diag (RWKV-7, rank off) | 0.390 [0.251, 0.691] | diagonal (vector decay) |
| mamba2 ref noconv | 0.151 [0.079, 0.191] | diagonal + selectivity, no conv |
Axis 1: The convolution (≈ +0.45, family-transferable)
- Delta-rule cell: 0.518 → 0.983 (+0.47; gated deltanet → armed gdn)
- Diagonal cell: 0.151 → 0.592 (+0.44; mamba2 ref noconv → mamba2 ref)
- The convolution is a state-free local primitive, yet the largest single lever in the factorial.
Axis 2: The transition (≈ +0.3 without convolution; ≈ +0.03 with it)
- Without convolution: dg rank vs. dg diag = 0.653 vs. 0.390 (+0.26); hardened: +0.191 at K=16 (p=0.0002), +0.319 at K=32 (p<10⁻⁴)
- With convolution: 0.992 (rank-1) vs. 0.958 (diagonal), +0.034 (p=0.0023) at K=32
Axis 3: Decay (no measurable cost)
- deltanet (no decay) 0.561 vs. gated deltanet (decay) 0.518 at K=32: p=0.86 at n=20
- Decay appears to trade occasional lucky solves for consistency; neither is a mean-level effect.
The Hardened Transition Ablation
Table 3: Hardened transition ablation (n=10 seeds)
| K | Arm | Mean [95% CI] | Lock-in ≥ 0.9 | Rebound | Drop |
|---|---|---|---|---|---|
| 16 | dg rank (rank-1) | 0.998 [0.995, 0.999] | 10/10 | 0.024 | +0.974 |
| 16 | dg diag (diagonal ablation) | 0.806 [0.721, 0.892] | 3/10 | 0.038 | +0.769 |
| 16 | deltanet (rank-1) | 1.000 [0.999, 1.000] | 10/10 | 0.024 | +0.976 |
| 16 | mamba2 ds16 (diagonal, matched) | 0.999 [0.999, 1.000] | 10/10 | 0.024 | +0.975 |
| 32 | mamba2 ds16 (diagonal, matched) | 0.849 [0.807, 0.890] | 3/10 | 0.042 | +0.807 |
| 32 | dg rank (rank-1) | 0.702 [0.631, 0.769] | 0/10 | 0.051 | +0.651 |
| 32 | deltanet (rank-1) | 0.623 [0.482, 0.741] | 1/10 | 0.057 | +0.566 |
| 32 | dg diag (diagonal ablation) | 0.382 [0.322, 0.450] | 0/10 | 0.073 | +0.309 |
Key findings:
- The within-cell ablation replicates decisively: dg rank > dg diag
- No class claim survives: A state- and parameter-matched Mamba-2 ties the rank-1 cell at K=16 (p=0.25) and beats it at K=32 (0.849 vs. 0.702, p=0.0036)
- The two statements together resolve the folklore: convolution-free recurrent cells lose to Mamba because of the convolution, not the recurrence
The Interference Wall
Table 4: The haystack wall, flat in length (n=5 seeds per cell)
| Arm (2 layers) | L=64 | L=128 | L=256 |
|---|---|---|---|
| attention (head dim 16) | 1.000 | 1.000 | 1.000 |
| armed gdn | 0.018 | 0.018 | 0.019 |
| mamba2 ref | 0.017 | 0.019 | 0.019 |
| dg rank (RWKV-7) | 0.018 | 0.018 | 0.018 |
Four discriminating properties:
- Not capacity: Same cells store 32 competing bindings at ≥0.99 but fail on 4 bindings plus distractors
- A wall, not a dip: Failure is total at the shortest tested length, inside the training distribution, flat through 4× the length
- Transition richness does not help: GDN's delta rule, Mamba's selectivity, and RWKV-7's removal key wall identically
- Attention is untouched: Per-token KV entries cannot be overwritten by later writes
The Curriculum Fix
Table 5: Curriculum lever attribution (n=10 seeds, L=64 haystack, 2 layers)
| Cell | Dense | Curric. | Lock-in | Per-seed |
|---|---|---|---|---|
| armed gdn (bidir) | 4 | ✓ | 6/10 | 1.00/0.87/0.88/0.99/0.36/0.45/0.12/0.04/1.00/0.97 |
| armed gdn (bidir) | 1 | ✓ | 7/10 | 0.82/0.13/0.06/1.00/0.92/1.00/0.58/1.00/1.00/1.00 |
| armed gdn (bidir) | 4 | — | 1/10 | 0.02/0.02/0.02/1.00/0.02/0.02/0.02/0.02/0.02/0.02 |
| armed gdn fwd (causal) | 4 | ✓ | 5/10 | 0.84/0.75/0.96/1.00/0.06/0.83/0.62/0.24/0.08/0.96 |
| attention (abs. pos.) | 4 | ✓ | 0/5 | all ≈ 0.046 |
| attention (rotary) | 4 | ✓ | 6/10 | mean 0.84; range 0.28–1.00 |
Key results:
- The distance curriculum lifts the walled cell from 0.021 (chance) to 1.000—an existence proof by construction
- Lock-in rises from 1/10 seeds (dense only) to 7/10 (curriculum), p=0.02 (Fisher two-sided)
- Adding dense supervision to the curriculum changes nothing (6/10)
- Length boundary: At L=256, uniform curriculum fails everywhere (0/9 trials) but a shaped ramp locks in 4/5 seeds (p≈0.005); at L=512, the time-based ramp collapses (0/6) while a success-gated ramp locks in 6/6 (p=0.0011)
Bidirectional Denoisers
Table 6: Collision-key discriminator (L=64, dense-4 + curriculum, 2 layers, n=10 seeds)
| Arm | Lock-in | Mean | Per-seed range |
|---|---|---|---|
| armed gdn (bidirectional) | 3/10 | 0.383 | 0.023–1.000 |
| armed gdn fwd (causal) | 1/10 | 0.301 | 0.036–0.844 |
| gated deltanet (bidir, no conv) | 0/10 | 0.128 | 0.017–0.238 |
Key findings:
- No directional margin at ten seeds (mean difference +0.082, p=0.63)
- Both directions can solve the task; any solving circuit needs two layers (mark-and-route)
- The convolution-free cell never locks in even under the shaped curriculum (0/10 vs. 9/10, p≈10⁻⁴)
- Curriculum shape is a powerful direction-agnostic stabilizer (1/10 → 9/10, p≈6×10⁻⁴)
Guardrail: Arming Is Free on State Tracking
Table 7: S₅ guardrail (d=128; chance ≈ 0.008)
| S₅ depth | armed gdn | gated deltanet | p |
|---|---|---|---|
| L=2 (n=10) | 1.000 [1.000, 1.000] | 0.980 [0.945, 1.000] | 0.0002 |
| L=4 (n=10) | 0.524 [0.492, 0.637] | 0.389 [0.274, 0.514] | 0.0014 |
| L=8 (n=10) | 0.214 [0.143, 0.261] | 0.152 [0.132, 0.240] | 0.0044 |
| L=16 (n=10) | 0.087 [0.072, 0.128] | 0.071 [0.070, 0.073] | <10⁻⁴ |
The armed cell is significantly better at every probed depth: arming for recall does not trade away state tracking.
Theoretical and Practical Implications
Re-scoping "Recurrent Models Are Bad at Recall"
At matched state, the claim decomposes into five things:
- A missing convolution (large, fixable for free in state terms)
- A transition effect (real, within-cell)
- A decay tax (real, scalar-gate-specific)
- An interference failure under sparse supervision (fixable by curriculum, within a length boundary)
- A directionality property (bidirectional denoisers get query-first reading natively)
None of these is "the recurrence cannot bind."
When Is a Hybrid Actually Needed?
Two classic triggers for "add attention layers for retrieval" dissolve under cheap interventions:
- The unarmed-cell recall deficit (add the convolution)
- The within-capacity haystack wall (train with a distance curriculum)
Three things remain as genuine hybrid territory:
- Loads beyond state capacity (the graceful K-limit)
- Lengths beyond the curriculum boundary (L ≥ 256 here, pending shaped curricula)
- Query-conditioned regimes a fixed-state writer cannot anticipate
For Recurrent Diffusion LMs
The JRT-for-free result gives bidirectional recurrent denoisers a principled recall story that causal recurrent LMs lack: the objective itself buys the second read. The requirements are concrete—two layers minimum and a short convolution in the backward stream.
The Central Conjecture
Within the capacity regime of a fixed-state recurrence, apparent retrieval walls are training-coverage gaps.
The length boundary at L=512 is its first direct test: the cell that sits at 0/6 under a time-based ramp locks in 6/6 once the ramp is gated on measured competence, so that boundary was the schedule, not the circuit.
Conclusion
Main Takeaways
-
The convolution is the recall story at constrained state: An armed diagonal cell reaches 0.958, showing that at constrained state the short convolution is very nearly the whole recall story.
-
No architecture-class claim survives matched-state comparison: The rank-1 transition's advantage over diagonal ablations is real but within-cell; a state-matched Mamba-2 ties the rank-1 cell, so "rank-1 > diagonal" as a class claim is false.
-
The interference wall is a training-coverage gap: The same architecture that sits at chance under naive training reaches 1.000 under a distance curriculum—an existence proof that recurrent retrieval failure is substantially a learnability problem.
-
Training is a lock-in lottery: Seed outcomes are bimodal (a seed either locks in or does not), and curriculum shape is the lever that controls the rate. Success-gated ramps reopen boundaries that time-based schedules cannot.
-
Arming for recall is free: The convolution-equipped cell is significantly better on S₅ state tracking at every depth, showing no trade-off between recall capability and state-tracking learnability.
Future Directions
- Whether some length defeats the gated schedule too (testing the conjecture's boundary)
- Whether the query-first mechanism carries to scale (untested beyond toy scale)
- Whether context-modulated forgetting (Kimi Team, 2025) becomes a useful design axis with better-powered tests
- The companion paper (Boesch & Wee, 2026) takes the curriculum half to converted diffusion LMs at 1.7B and 8B parameters
Limitations
- Toy scale (, synthetic tasks)
- The Mamba-2 reference under-trains the official kernel
- Learning-rate sensitivity affects magnitude claims
- The attention control is valid only where reported (two-layer regimes)
- Cross-family magnitude comparisons carry a parameterization caveat
Related papers
- Fast Learning Rate Transfer in Shallow Linear Networks at Growing Training Horizons
Fast learning-rate transfer holds at growing training horizons when T grows slower than sqrt(n), but requires nondegenerate first-order loss sensitivity to avoid spectral-dependent failures.
- Just Let Linear States Forget the Distant Past: Prefix Caching via Suffix Replay for Hybrid LLMs
SuffixReplay enables fine-grained prefix caching in hybrid LLMs by reconstructing linear-attention states from sparse anchors, cutting storage by 2x and TTFT by up to 70%.
- On Trajectory-Aware Training for Masked Diffusion Language Models
PUMBA trains masked diffusion language models on inference-like trajectories via continuous hidden-state carries and backpropagation through time, matching autoregressive accuracy while decoding multiple tokens per step.