Skip to main content
Back to timeline
arXivSource publication:

VoxPolyMem pairs interaction-aware hierarchical memory with an EG-GRPO retrieval policy to score 85.0 on the multi-party spoken-memory benchmark VoxPolyBench, 23.6 points above the strongest baseline

Synopsis

The work proposes VoxPolyMem, an interaction-aware multimodal long-term memory framework for multi-party spoken conversations that combines incremental speaker identification with a memory hierarchy of interaction memory, fact memory, and participant profiles, formulates retrieval as sequential decision-making, and trains it with Evidence-Gain GRPO (EG-GRPO) to reward newly acquired supporting evidence round by round; it also builds VoxPolyBench (18 scenarios, 176 sessions, 18.9 hours of synthesized speech, 1,527 QA pairs), on which VoxPolyMem scores 85.0 overall, surpassing the strongest evaluated baseline by 23.6 points, and scores 89.6 and 74.4 on Mem-Gallery and H2HMem-Multi, exceeding the strongest public memory baselines by more than 8 points each.

AI-generated editorial illustration: Beyond Dyadic Memory: Interaction-Aware Multimodal Memory with Adaptive Agentic Retrieval for Multi-Party Spoken Conversations

Interpretation

It introduces VoxPolyMem, which splits long-term memory of multi-party speech into interaction memory, fact memory, and participant profiles, and records who said what to whom in a directed interaction graph. Existing memory systems mostly organize conversational history through temporal graphs and semantic abstraction without explicitly storing interaction roles; speech-side work such as AFA uses voiceprints to separate users' memories but does not represent a shared multi-party dialogue. The paper defines the interaction graph's edge set (speaker and addressee sets) and the jointly extracted fields of facts and profiles, and reports an ablation in which removing interaction relations drops the VoxPolyBench score from 84.0 to 78.6.

It formulates retrieval as sequential decision-making, where a retrieval agent rewrites queries and selects memory layers and tools based on accumulated evidence, trained with EG-GRPO, which assigns credit per round for the coverage and ranking of newly acquired supporting evidence. Search-R1 and similar methods use terminal answer rewards, and Mem-T's MoT-GRPO combines tree-structured rollouts with evidence coverage and answer rewards; EG-GRPO computes rewards after each action and after context selection, excluding previously retained evidence. Under fixed memory, tools, and inference budget, EG-GRPO leads all three benchmarks in answer score and both annotated benchmarks in recall; on Mem-Gallery it raises answer score from 87.1 to 89.6 and recall from 91.5 to 93.6 over MoT-GRPO, while averaging 1.2–1.3 retrieval rounds versus 1.4–1.7 for the w/o RL variant.

It constructs VoxPolyBench, covering cross-session speaker identification, memory evolution, personalized answering, retrieval and reasoning, and interaction attribution, with gold supporting-evidence annotations. The paper's Table 1 comparison shows that speech benchmarks such as ContextDialog, MSU-Bench, AFA, and Vox-Infinity, and image-text memory benchmarks such as Mem-Gallery and H2HMem, do not jointly cover multi-party, multi-session, cross-session identity, interaction attribution, personalization, and memory update. The benchmark contains 18 scenarios, 176 sessions, 9,599 dialogue turns, 18.9 hours of synthesized speech, and 1,527 QA pairs; after synthesis, Whisper large-v3-turbo evaluation yields a corpus-level WER of 2.14%, alongside human review of event anchors and QA validation.

VoxPolyMem leads on VoxPolyBench and on two public image-text memory benchmarks. The paper reports a weighted average of 86.6 versus 68.1 for the strongest external baseline, and notes that the variant without RL already exceeds all external baselines with a weighted average of 84.4. VoxPolyBench overall score 85.0 (23.6 points above the strongest external baseline), Mem-Gallery 89.6 (+11.8), H2HMem-Multi 74.4 (+8.4); all methods and ablations share GPT-4.1-mini as the answer model, shared questions and rubric, and two judges (GPT-4.1-mini and GPT-4.1); in a 200-response human audit, aggregate LLM and human means were 84.6 and 86.1.

Perspective

The result targets multi-party voice assistant settings that must remember participant identities and interaction relations across sessions, such as meetings, in-car, and household assistance; methodologically, the hierarchical memory and the EG-GRPO reward design can transfer to other multimodal long-term memory tasks, and the paper also reports leading results on the two public image-text memory benchmarks Mem-Gallery and H2HMem-Multi. For a reader, it offers a reusable idea: store who said what to whom as a first-class part of memory, and let the retrieval policy pay per round for new evidence rather than only for the final answer.

VoxPolyBench uses generated dialogues and synthesized speech, and the ethics statement states that this controlled setting "does not establish robustness across real speakers, accents, or recording conditions," so behavior under real accents and recording conditions remains open; the speaker-tracking evaluation reuses frozen ECAPA embeddings and reference utterance boundaries, isolating identity tracking rather than end-to-end diarization; EG-GRPO training depends on supporting-turn annotations that are unavailable at inference, where stopping instead uses evidence sufficiency and the round budget; and although the human audit shows close aggregate LLM and human means, the paper notes that "Similar means do not establish item-level agreement," leaving per-item agreement to be examined. This reading covered the full text, but figures and tables were rendered as text and their visual details were not checked item by item.

Sources