Summary

Overview

  • Introduces VoxMem, a benchmark for evaluating multimodal memory in Large Audio Language Models (LALMs), spanning 3,196 evaluation instances over 34,743 spoken sessions (177 hours).
  • Proposes a two-dimensional taxonomy jointly characterizing acoustic evidence (speech semantics, speaker identity, paralinguistic cues, environmental sound) and memory operations (information extraction, multi-session reasoning, temporal tracking, answer refusal).
  • Evaluates 15 LALMs and finds that models retain what was said far better than who said it, how, or what was audible, with gaps widening for complex operations and longer histories.
  • Provides fine-grained error analysis showing qualitatively distinct failure modes across evidence types, indicating distinct underlying deficiencies rather than a shared bottleneck.
  • All questions are designed to be unanswerable from transcripts alone for audio-native items, verified via transcript-only baselines.

Introduction and Theoretical Foundation

  • Background: LALMs increasingly support long, multi-session spoken interactions, but memory—the ability to accumulate and use information across past interactions—remains under-evaluated.
  • Problem: Existing benchmarks focus on lexical content, use ad hoc memory operations, and treat memory as a single-session problem. They fail to evaluate memory for who spoke, how they spoke, or what was audible.
  • Proposed Framework: A two-axis taxonomy:
    • Acoustic evidence: What must be remembered from speech (semantics, speaker identity, paralinguistic cues, environmental sound).
    • Memory operations: How the evidence must be used (extraction, multi-session reasoning, temporal tracking, answer refusal).
  • Key Insight: Memory evaluation must jointly consider the type of information to be retained and the operation applied to it, especially across distinct conversational episodes.

Methodology

  • Benchmark Construction:
    • Built from multi-session histories with evidence, haystack, and filler sessions across 20 topic families.
    • Questions are open-ended with short answers; audio-native items require acoustic evidence not present in transcripts.
    • 15 evaluation scenarios (4 evidence types × 4 operations, minus one invalid combination).
  • Quality Control:
    • QC-1: Audio-language model verifies transcript fidelity, speaker consistency, and perceptibility of acoustic cues.
    • QC-2: Gemini-3.7-Flash filters questions answerable from audio but not transcript; transcript-only accuracy drops from 46.2% to 4.4% on audio-native questions.
    • QC-3: Text-only classifiers cannot distinguish evidence from haystack sessions (51–56% accuracy vs. 47–49% baseline).
  • Evaluation Setup:
    • Models evaluated via native audio interfaces.
    • Context lengths: 8K, 16K, 32K, 64K tokens.
    • Question–evidence pairs held fixed across context budgets to isolate the effect of history length.

Empirical Validation / Results

  • Overall Performance: No model exceeds 40% overall accuracy at 32K, indicating that multi-session spoken memory remains unsolved.
  • Evidence Type Difficulty:
    • Speech semantics: ~75.9% (transcript-only: 71.0%)
    • Speaker identity: ~20.0% (temporal tracking)
    • Paralinguistic: ~3.4% (temporal tracking)
    • Environmental: ~1.2% (temporal tracking)
  • Memory Operations:
    • Multi-session reasoning is relatively robust (e.g., 41.8% on semantics, 43.1% on speaker identity).
    • Temporal tracking collapses for non-lexical evidence, suggesting weak stateful modeling of acoustic cues.
  • Error Analysis:
    • Speaker identity errors: 48% are binding/association failures (wrong attribution).
    • Paralinguistic/environmental errors: More often reflect failure to retain the acoustic cue itself.
    • Information extraction errors: 46% are binding failures, indicating difficulty linking cues to context.

Theoretical and Practical Implications

  • Theoretical:
    • Memory for non-lexical acoustic information is not a homogeneous challenge; different evidence types exhibit distinct failure modes.
    • Larger context windows do not guarantee stable access to information; accuracy drops with longer histories even when the answer is present.
    • Transcript-only evaluations overestimate LALM memory capabilities and miss audio-native failures.
  • Practical:
    • Future systems must preserve non-lexical acoustic information and maintain associations with speakers, events, and sessions over time.
    • Benchmarks must jointly evaluate evidence type and operation to guide development of robust spoken conversational memory.

Conclusion

  • VoxMem provides the first principled, multi-session benchmark for spoken conversational memory, spanning 15 scenarios across 4 evidence types and 4 operations.
  • Findings show that LALMs retain lexical content well but struggle with speaker identity, paralinguistic cues, and environmental sound, especially under temporal tracking and long histories.
  • The benchmark enables systematic evaluation and drives progress toward more robust, memory-capable spoken dialogue systems.
  • Future Directions: Improve stateful modeling of acoustic cues, strengthen cross-session binding, and develop models that can abstain appropriately when evidence is absent.

"Models retain what was said far better than who said it, how, or what was audible... Future systems must preserve non-lexical acoustic information and maintain its associations with speakers, events, and sessions over time."

Related papers