VoxMem tests 15 audio LLMs on 3,196 questions: none tops 40% at 32K, and swapping audio for transcripts drops speaker accuracy from 69.8% to 10.3%
Synopsis
The authors propose a two-axis taxonomy of spoken conversational memory — acoustic evidence type (speech semantics, speaker identity, paralinguistic cues, environmental sound) crossed with memory operation (information extraction, multi-session reasoning, temporal evolution tracking, answer refusal) — and build VoxMem on it: 3,196 evaluation instances over 34,743 spoken sessions (177 hours) at four context budgets from 8K to 64K tokens; evaluating 15 large audio language models, no model exceeds 40% overall accuracy at 32K (best 38.5%), models remember what was said far better than who said it, how it was said, or what was audible, and replacing audio with exact transcripts drops speaker accuracy from 69.8% to 10.3% while speech-semantics accuracy barely moves (75.9% to 71.0%).
Interpretation
The paper introduces a two-axis taxonomy that jointly characterizes the acoustic evidence to be remembered and the memory operations applied to it, and builds VoxMem on it: 3,196 evaluation instances over 34,743 spoken sessions (177 hours), spanning 15 evaluation scenarios and four context budgets (8K, 16K, 32K, 64K). Existing spoken benchmarks focus primarily on lexical content, adopt limited and ad hoc memory operations, and treat memory as a single-session problem; VoxMem is, to the authors' knowledge, the first spoken benchmark built on multi-session histories composed of evidence, competing haystack, and unrelated filler sessions across 20 topic families, with each question held fixed across context lengths. Benchmark statistics are given in the paper's tables (799 questions, 3,196 instances, 34,743 sessions, 1,326 evidence sessions, about 9 turns per session, 2.5 to 20 minutes of audio per instance); construction passes two levels of quality control, and the acoustic-necessity check uses Gemini-3.7-Flash, which is not among the evaluated models.
Across 15 large audio language models, no model exceeds 40% overall accuracy at the 32K budget, with the strongest reaching 38.5%; the five proprietary models average 33.0% and the ten open-weight models 21.9%. Previously there was no unified multi-session, multi-evidence, multi-operation characterization of LALM memory; this result turns 'longer context is not reliable memory' from an assumption into a measurable phenomenon. Fifteen models (ten open-weight, five proprietary) are evaluated through their native audio interfaces, with length measured on a shared Whisper-encoder token scale; answers are scored by Gemini-3.7-Flash under answer-type rules and re-scored by GPT-5.6-Luna, with 97.95% agreement on 11,978 8K predictions.
Memory for non-lexical acoustic information is substantially weaker than for speech semantics: at 32K, proprietary models average 55.6% on speech semantics versus 32.7% on speaker identity, 20.0% on paralinguistic cues, and 21.9% on environmental sound; open-weight models average 31.6%, 26.7%, 14.5%, and 15.5% respectively. The result separates audio-native information from transcript-recoverable semantics, showing that strong lexical performance masks broader failures to preserve what spoken interaction conveys. The paper reports a full-audio versus transcript-only comparison: on 669 answerable questions, transcript-only input reduces accuracy on audio-native questions from 46.2% to 4.4%, while speech semantics only falls from 75.9% to 71.0%; the community summary further reports speaker accuracy falling from 69.8% to 10.3%.
Memory difficulty depends on the interaction between operation and evidence type: multi-session reasoning is relatively robust (41.8% on speech semantics, 43.1% on speaker identity), while temporal evolution tracking reaches 44.5% on speech semantics but collapses to 3.4% on paralinguistic cues and 1.2% on environmental sound; answer refusal shows the opposite profile, higher on paralinguistic and environmental questions (34.3%, 34.2%) than on speech semantics and speaker identity (19.9%, 16.2%). The paper argues that high refusal accuracy on acoustic questions more likely reflects a general inability to identify or use acoustic evidence than genuine awareness that evidence is insufficient, separating refusal from answering as distinct capabilities. The paper reports accuracy stratified by evidence type and operation, plus error attribution: 48% of speaker errors are binding failures, 63% of paralinguistic errors are localization failures, and environmental errors split between localization (41%) and unsupported answers (39%).
Perspective
The work targets researchers and system developers evaluating multi-session spoken conversational memory: the taxonomy and benchmark can compare LALM acoustic memory at 8K to 64K context budgets and separate differences by evidence type and operation. Histories are built from synthesized speech (Higgs-TTS-3 user voices, fixed VCTK references, ESC-50 environmental sounds mixed at 10 dB SNR), and assistant turns are provided as text, so results apply to audio-input models under this setting; the paper also notes that length is measured on a Whisper-encoder scale that does not represent how many tokens any evaluated model spends on an item.
Evaluation relies on an LLM judge (Gemini-3.7-Flash) for open-ended short answers; the paper reports 97.95% agreement with GPT-5.6-Luna and notes the production judge favors Gemini outputs by about 0.88 percentage points, and that at 8K the three strongest models lie within 0.9 percentage points, so their order at that budget should not be read as meaningful. The 64K results cover only the 10 of 15 models that completed the full 64K run. The error analysis uses the 64K budget, the three audio-native evidence types, and those 10 models, excluding refusals, empty generations, and truncated answers. Histories are synthesized speech, so whether real recordings with accents, noise, and overlapping speech show the same patterns remains an open question.
