# VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models

> VoxMem reveals large audio-language models remember speech content far better than speaker identity, paralinguistic cues, or environmental sounds, with accuracy collapsing under temporal tracking.

- **Source:** [arXiv](https://arxiv.org/abs/2609.32607)
- **Published:** 2026-10-01
- **Permalink:** https://picx.dev/p/l9bWKH
- **Whiteboard:** https://picx.dev/p/l9bWKH/image

## Summary

## Summary

### Overview
- Introduces **VoxMem**, a benchmark for evaluating multimodal memory in Large Audio Language Models (LALMs), spanning 3,196 evaluation instances over 34,743 spoken sessions (177 hours).
- Proposes a **two-dimensional taxonomy** jointly characterizing acoustic evidence (speech semantics, speaker identity, paralinguistic cues, environmental sound) and memory operations (information extraction, multi-session reasoning, temporal tracking, answer refusal).
- Evaluates **15 LALMs** and finds that models retain *what was said* far better than *who said it*, *how*, or *what was audible*, with gaps widening for complex operations and longer histories.
- Provides fine-grained error analysis showing **qualitatively distinct failure modes** across evidence types, indicating distinct underlying deficiencies rather than a shared bottleneck.
- All questions are designed to be **unanswerable from transcripts alone** for audio-native items, verified via transcript-only baselines.

---

### Introduction and Theoretical Foundation
- **Background**: LALMs increasingly support long, multi-session spoken interactions, but memory—the ability to accumulate and use information across past interactions—remains under-evaluated.
- **Problem**: Existing benchmarks focus on lexical content, use ad hoc memory operations, and treat memory as a single-session problem. They fail to evaluate memory for *who* spoke, *how* they spoke, or *what was audible*.
- **Proposed Framework**: A two-axis taxonomy:
  - **Acoustic evidence**: What must be remembered from speech (semantics, speaker identity, paralinguistic cues, environmental sound).
  - **Memory operations**: How the evidence must be used (extraction, multi-session reasoning, temporal tracking, answer refusal).
- **Key Insight**: Memory evaluation must jointly consider the *type* of information to be retained and the *operation* applied to it, especially across distinct conversational episodes.

---

### Methodology
- **Benchmark Construction**:
  - Built from multi-session histories with **evidence, haystack, and filler sessions** across 20 topic families.
  - Questions are open-ended with short answers; audio-native items require acoustic evidence not present in transcripts.
  - 15 evaluation scenarios (4 evidence types × 4 operations, minus one invalid combination).
- **Quality Control**:
  - **QC-1**: Audio-language model verifies transcript fidelity, speaker consistency, and perceptibility of acoustic cues.
  - **QC-2**: Gemini-3.7-Flash filters questions answerable from audio but not transcript; transcript-only accuracy drops from 46.2% to 4.4% on audio-native questions.
  - **QC-3**: Text-only classifiers cannot distinguish evidence from haystack sessions (51–56% accuracy vs. 47–49% baseline).
- **Evaluation Setup**:
  - Models evaluated via native audio interfaces.
  - Context lengths: 8K, 16K, 32K, 64K tokens.
  - Question–evidence pairs held fixed across context budgets to isolate the effect of history length.

---

### Empirical Validation / Results
- **Overall Performance**: No model exceeds **40% overall accuracy at 32K**, indicating that multi-session spoken memory remains unsolved.
- **Evidence Type Difficulty**:
  - Speech semantics: ~75.9% (transcript-only: 71.0%)
  - Speaker identity: ~20.0% (temporal tracking)
  - Paralinguistic: ~3.4% (temporal tracking)
  - Environmental: ~1.2% (temporal tracking)
- **Memory Operations**:
  - Multi-session reasoning is relatively robust (e.g., 41.8% on semantics, 43.1% on speaker identity).
  - Temporal tracking collapses for non-lexical evidence, suggesting weak stateful modeling of acoustic cues.
- **Error Analysis**:
  - **Speaker identity errors**: 48% are binding/association failures (wrong attribution).
  - **Paralinguistic/environmental errors**: More often reflect failure to retain the acoustic cue itself.
  - **Information extraction errors**: 46% are binding failures, indicating difficulty linking cues to context.

---

### Theoretical and Practical Implications
- **Theoretical**:
  - Memory for non-lexical acoustic information is not a homogeneous challenge; different evidence types exhibit distinct failure modes.
  - Larger context windows do **not** guarantee stable access to information; accuracy drops with longer histories even when the answer is present.
  - Transcript-only evaluations overestimate LALM memory capabilities and miss audio-native failures.
- **Practical**:
  - Future systems must preserve non-lexical acoustic information and maintain associations with speakers, events, and sessions over time.
  - Benchmarks must jointly evaluate evidence type and operation to guide development of robust spoken conversational memory.

---

### Conclusion
- VoxMem provides the first principled, multi-session benchmark for spoken conversational memory, spanning 15 scenarios across 4 evidence types and 4 operations.
- Findings show that LALMs retain lexical content well but struggle with speaker identity, paralinguistic cues, and environmental sound, especially under temporal tracking and long histories.
- The benchmark enables systematic evaluation and drives progress toward more robust, memory-capable spoken dialogue systems.
- **Future Directions**: Improve stateful modeling of acoustic cues, strengthen cross-session binding, and develop models that can abstain appropriately when evidence is absent.

---

> *"Models retain what was said far better than who said it, how, or what was audible... Future systems must preserve non-lexical acoustic information and maintain its associations with speakers, events, and sessions over time."*

---

_Markdown view of https://picx.dev/p/l9bWKH, served by PicX — AI-generated visual whiteboard summaries of research papers._
