# Balancing Memory Pathways: Analyzing and Improving Memory Utilization in Hybrid LMs

> Hybrid language models over-rely on attention and underuse recurrent memory, but an auxiliary loss forcing recurrent routing improves QA accuracy by up to 5.2% and agentic success rates.

- **Source:** [arXiv](https://arxiv.org/abs/2610.06750)
- **Published:** 2026-10-10
- **Permalink:** https://picx.dev/p/HNnZ3U
- **Whiteboard:** https://picx.dev/p/HNnZ3U/image

## Summary

# Balancing Memory Pathways: Analyzing and Improving Memory Utilization in Hybrid LMs

## Summary (Overview)

- **Core Problem**: Recurrent-attention hybrid language models (LMs) contain two memory pathways—attention and recurrent layers—but simply having access to both does not ensure models use both effectively. The authors find a substantial over-reliance on attention, with recurrent pathways remaining underutilized.

- **Key Finding**: Standard supervised fine-tuning(SFT) improves overall performance but exacerbates the imbalance by increasing reliance on attention while leaving recurrent pathway usage largely unchanged.



- **Proposed Solution**: The authors introduce an auxiliary loss($\mathcal{L}_{\text{rec}}$) that masks attention's access to past context during an auxiliary forward pass, forcing the model to route information through recurrent layers. This encourages better coordination between memory pathways.





- **Results**: Adding the auxiliary loss improves average QA accuracy by 3.9% on Qwen3.5-4B and 5.2% on Nemotron-H-4B-Instruct across six datasets, with particularly large gains on longer-context tasks (8.0% improvement vs. 0.2% on shorter contexts). The approach also improves agentic task success rates and recovery from past failures.



- **Generalization**: The imbalance and benefits of auxiliary supervision extend to multiple recurrent-attention LMs, question-answering and agentic tasks, and even to attention-based LMs combining different forms of memory (textual and latent).



## Introduction and Theoretical Foundation

### Background and Motivation

Recurrent-attention hybrid LMs interleave attention layers and linear recurrent layers(such as Mamba2 or Gated DeltaNet) to combine computational efficiency with strong performance. These architectures contain two distinct memory pathways:

- **Attention layers**: Provide direct access to earlier tokens, supporting precise memory recall at the cost of quadratic computation in sequence length
- **Recurrent layers**: Update a fixed-size state as tokens are processed, enabling linear computation and constant memory, supporting consolidation of information distributed across long contexts

Prior work(Lee et al.,, 2025; Michalak & Abreu, 2025; Bick et al.,, 2026) has shown these pathways have complementary strengths:

- Attention is critical for **precise retrieval**—ablating attention causes severe retrieval failures
- Recurrent layers contribute substantially to **language modeling**—ablating them causes larger degradations in likelihood-based evaluation than removing attention

The authors extend these findings by analyzing performance across question types:

- Attention plays a larger role when answers appear **verbatim** in context or require **single evidence**
- Recurrent state plays a larger role when answers must be **inferred** or require **aggregating multiple pieces of evidence**

This complementary relationship motivates the central research question: *How well do hybrid LMs use both memory pathways, and how can their coordination be improved?*

### Theoretical Foundation

The paper builds on the observation that attention and recurrent layers offer complementary memory pathways. The authors hypothesize that effective coordination of both pathways should leverage their respective strengths—attention for precise recall and recurrence for information aggregation over long contexts. However, they demonstrate that hybrid LMs fail to achieve this balance naturally or through standard fine-tuning, motivating targeted supervision to encourage recurrent pathway usage.



## Methodology

### Models

The authors experiment with two recurrent-attention hybrid LMs:

| Model | Recurrent Layers | Attention Layers |
|-------|------------------------|---------------------|
| Qwen3.5-4B | 24 Gated DeltaNet | 8 full-attention |
| Nemotron-H-4B-Instruct | 24 Mamba2 | 4 full-attention |

### Training Setup

- **Method**: LoRA on all layers with rank 16, scaling factor 32, dropout 0.05
- **Optimizer**: AdamW with learning rate 1e⁻⁴, cosine schedule, warmup covering 3% of steps, effective batch size 16
- **Compute matching**: Objectives with auxiliary loss use half the optimization steps to match total training compute with standard SFT

### Disentangling Memory Pathways

To analyze how past information flows through each pathway, the authors define two complementary interventions:

1. **Recurrent-only condition**: Attention is blocked across segment boundaries while recurrent state propagates normally. The attention mask is defined as:

$$
M _ {t, s} = \left\{ \begin{array}{l l} 0, & s \leq t \text {   and   } \sigma (s) = \sigma (t), \\ - \infty , & \text { otherwise }. \end{array} \right.
$$

where $M_{t,s}$ is added to the attention score between query position $t$ and key position $s$ before normalization, and $\sigma(t)$ denotes the segment containing token $t$.



2. **Attention-only condition**: Standard causal attention is retained, but recurrent state propagation across segment boundaries is blocked:

$$

S _ {t} = \mathbb {1} [ \sigma (t) = \sigma (t - 1) ]   \mathcal {A} _ {t} (S _ {t - 1}) + w _ {t}, \qquad S _ {0} = 0,

$$

where $S_t$ denotes the recurrent state after processing token $t$, $\mathcal{A}_t$ denotes the recurrent transition, $w_t$ denotes the contribution of the current token, and $\mathbb{1}[\cdot]$ is an indicator functionthat zeros out the previous state at segment boundaries.



### Auxiliary Loss for Recurrent Supervision

The authors augment standard SFT with an auxiliary loss computed during a forward pass where attention is masked from accessing preceding context, while recurrent layers remain unrestricted:

$$

\mathcal {L} (\theta) = \mathcal {L} _ {\mathrm{SFT}} (\theta) + \lambda \mathcal {L} _ {\mathrm{rec}} (\theta),
$$

where $\theta$ denotes model parameters and $\lambda$ controls the strength of the auxiliary supervision. This forces information from preceding context to reach the current portion only through recurrent memory, encouraging the model to retain and use information via the recurrent pathway.



### Tasks and Evaluation

- **Question Answering**: Training on DROP, HotpotQA, Qasper, NarrativeQA(12k training examples). Evaluation additionally on MINTEval and MuSR(8.2k test examples, average context length 84k tokens)
- **Agentic Tasks**: TextWorld(Quest, Treasure)and BabyAI(PutNextLocal, GoToObjMaze), where agents must explore, accumulate observations, and choose actions accordingly



## Empirical Validation / Results

### Standard SFT Favors Attention Over Recurrence

Before SFT, Qwen3.5 achieves 30.2% QA accuracy under full conditions. Under attention-only conditions, accuracy drops to 20.2%(67% retention), while recurrent-only conditions cause a severe drop to 5.1%(17% retention). This indicates strong reliance on attention.



After standard SFT:

- Full accuracy improves to 34.6%
- Attention-only accuracy improves to 25.0%
- Recurrent-only accuracy remains nearly unchanged(5.0%

Similar patterns appear in agentic Quest task: full success rises from 27.0% to 78.5%, attention-only from 6.0% to 52.0%, but recurrent-only improves only slightly(from 2.5% to 4.0%).

Attention-only retention increases for all question types after SFT, with larger gains for verbatim(+25.5%avg) and single-evidence questions compared to non-verbatim and multi-evidence(+19.2%.

### Auxiliary Loss Improves Performance

**Question Answering Results(Table 1):

| Model | Training Obj. | DROP | MuSR | HQA | Qasper | NQA | MINT | Avg. |
|-------|----------------|------|------|------|--------|------|-------|-------|
| Qwen3.5 | $\mathcal{L}_{\text{SFT}}$ | 48.6 | 5.6 | 3.1 | 0.1 |  ̄4.4 |  ̄1.1 |  ̄4.6 |
| | $\mathcal{L}_{\text{SFT}} + \mathcal{L}_{\text{rec}}$ (Ours) | 46..9 |  ̄4..8 |  ̄4..0 |  ̄5..9 |  ̄3..7 |  ̄ ̄8..9 |  ̄ ̄8..5 |
| Nemotron-H | $\mathcal{L}_{\text{SFT}}$ |  ̄ ̄2..4 |  ̄ ̄0..9 |  ̄ ̄8..7 |  ̄ ̄8..1 |  ̄ ̄1..9 |  ̄ ̄8..8 |  ̄ ̄1..0 |
| | $\mathcal{L}_{\text{SFT}} + \mathcal{L}_{\text{rec}}$ (Ours) |  ̄ ̄4..5 |  ̄ ̄8..1 |  ̄ ̄ ̄2..1 |  ̄ ̄ ̄4..1 |  ̄ ̄1..7 |  ̄ ̄ ̄7..8 |  ̄ ̄ ̄6..2 |

Adding recurrent auxiliary loss improves average accuracy from 34..6% to 38..5%(Qwen3.5)and from 31.0% to 36.2%(Nemotron-H).The largest gains appear on Qasper, NarrativeQA, and MINTEval—datasets with longer contexts.



Recurrent-only accuracy improves substantially from  ̄5.0% to 22..8%, while attention-only accuracy is largely preserved(25..0% to 23..7%), indicating more effective recurrent pathway usage.



**Ablation analysis** reveals:

- $\mathcal{L}_{\text{SFT}} + \mathcal{L}_{\text{attn}}$(encouraging attention)yields smaller gains than $\mathcal{L}_{\text{SFT}} + \mathcal{L}_{\text{rec}}$(2.1% vs. 3..9% for Qwen3.5; 4..6% vs. 5..2% for Nemotron-H)
- Training with only auxiliary losses yields lower accuracy than standard SFT, confirming restricted losses work best as auxiliary objectives alongside unrestricted SFT



**Agentic Task Results(Table 2):

| Model | Training Obj. | TextWorld Quest | TextWorld Treasure | BabyAI PutNextLocal | BabyAI GoToObjMaze |
|-------|----------------|------------------|---------------------|------------------------|----------------------|
| Qwen3.5 | $\mathcal{L}_{\text{SFT}}$ | 78..5 |  ̄4..0 |  ̄ ̄6..0 |  ̄ ̄3..0 |
| | $\mathcal{L}_{\text{SFT}} + \mathcal{L}_{\text{rec}}$ (Ours) |  ̄0..5 |  ̄ ̄6..0 |  ̄ ̄3..0 |  ̄ ̄3..0 |
| Nemotron-H | $\mathcal{L}_{\text{SFT}}$ |  ̄0..5 |  ̄ ̄0..5 |  ̄ ̄0..0 |  ̄ ̄ ̄7..0 |
| | $\mathcal{L}_{\text{SFT}} + \mathcal{L}_{\text{rec}}$ (Ours) |  ̄5..5 |  ̄ ̄ ̄7..0 |  ̄ ̄ ̄7..0 |  ̄ ̄ ̄4..0 |

Averaged across both models, adding recurrent auxiliary supervision improves success rates by:

- **TextWorld**: +3..5%(Quest)and +9.3%(Treasure)
- **BabyAI**: +22.0%(PutNextLocal)and +13.5%(GoToObjMaze)

The larger gains on BabyAI align with QA results—BabyAI trajectories average 134.0 steps versus 42.2 for TextWorld, consistent with recurrent layers' strength in longer contexts.



### Auxiliary Supervision Improves Learning from Past Trajectories

**Learning from Failure**: When models are allowed to retry after failures, with previous attempts stored in memory:

- QA recovery rate improves by 16.6% on average with the auxiliary loss
- Gains are especially large on multi-evidence questions(30.3% improvement vs. 2.8% for single-evidence
- Agentic tasks show 19.2% average improvement in cumulative recovery, whereas standard SFT plateaus with additional attempts

**Learning from Success**: When given multiple suboptimal successful demonstrations:

- Agents trained with the auxiliary loss generate trajectories shorter than any provided demonstration 6.0% more often on average
- Gains increase as demonstration length grows relative to optimal path



### Generalization to Attention-Based LMs

The imbalance generalizes to attention-based LMs with multiple memory types(textual and latent memory:

- Standard SFT with both memories achieves 24.1% accuracy, far below the Oracle upper bound(32.0%)and even below the text-only model(25.9%)
- Adding auxiliary supervision encouraging use of both pathways improves accuracy to 28.6%, driven by increased latent memory usage while preserving textual memory usage



## Theoretical and Practical Implications

### Theoretical Implications

1. **Access ≠ Effective Use**: The paper demonstrates that simply providing multiple memory pathways in an architecture does not ensure balanced, complementary use. Models develop strong biases toward one pathway(attention)even when recurrent layers offer complementary strengths.





2. **SFT Does Not Fix Imbalance**: Standard supervised fine-tuning improves overall task performance but does not inherently lead to better coordination of memory pathways—it can actually exacerbate existing biases by further favoring the already-dominant pathway.



3. **Targeted Supervision is Needed**: Deliberate training objectives that restrict access to one pathway during auxiliary training can effectively encourage use of the underutilized pathway, leading to better overall coordination and performance.



4. **Pathway Strengths Determine Optimal Supervision**: The improvement from auxiliary supervision depends on *which* pathway is encouraged—strengthening the underutilized recurrent pathway yields larger gains than further encouraging attention, consistent with the complementary strengths identified in analysis.



### Practical Implications

1. **Training Methodology**: The auxiliary loss approach provides a compute-matched method(halving optimization steps)to improve hybrid LM performance without requiring additional training compute.



2. **Long-Context Applications**: The approach is particularly beneficial for tasks requiring information aggregation over long contexts, such as long-document QA and long-horizon agentic tasks, where recurrent layers' strengths are most relevant.



3. **Agentic Systems**: Improved memory coordination translates to better recovery from failures and more efficient use of successful demonstrations, valuable for building robust agents that learn from experience.



4. **Architecture-Agnostic**: The findings apply beyond recurrent-attention hybrids to any LM architecture combining multiple memory pathways, suggesting broad applicability for improving memory utilization in diverse model designs.



## Conclusion

This paper reveals that recurrent-attention hybrid LMs exhibit a significant imbalance in memory pathway utilization, relying predominantly on attention while underusing recurrent layers. Standard SFT exacerbates this imbalance by further increasing attention reliance without improving recurrent pathway usage.



The authors propose a simple yet effective intervention: an auxiliary loss that masks attention's access to past context during training, forcing the model to route information through recurrent layers. This approach:

- Improves overall performance on QA tasks(+3.9% to 5.2% average accuracy)
- Yields particularly large gains on longer-context tasks
- Enhances agentic task success rates, especially in long-horizon environments
- Improves recovery from failures and efficient use of successful demonstrations
- Generalizes across architectures and memory types

The core message is that **providing multiple memory pathways does not ensure their effective use**—models must be explicitly trained to coordinate these pathways. The auxiliary supervision approach offers a practical, compute-matched method for achieving this coordination, with implications for designing and training more capable memory-augmented language models. Future work could explore extending this approach to other architectures, more diverse tasks, and adaptive strategies for balancing memory pathway usage dynamically during inference.

---

_Markdown view of https://picx.dev/p/HNnZ3U, served by PicX — AI-generated visual whiteboard summaries of research papers._
