Skip to main content
Back to timeline
arXivSource publication:

VGBench's 1,018 controlled items show voice agents recognize the tool but rarely stay silent when the speaker changes

Synopsis

The authors introduce VGBench, a 1,018-item diagnostic benchmark that tests whether audio LLMs condition execution on acoustic and conversational context across side-talk, self-talk, and speaker-switch scenarios under one silence/tool-call/natural-language action space; six raw audio LLMs and three training-free adaptations often identify the target tool yet rarely withhold action under the speaker switch (highest raw switch mute rate 14%), while a VoxGate post-training case study raises switch muting to 91.3% while keeping tool selection correct for nearby wearer commands and text-only controls.

AI-generated editorial illustration: Do Audio LLMs Listen Before They Act? Diagnosing Acoustic-Context Gating in Voice Agents

Interpretation

VGBench turns "should I act?" into an action-level diagnostic: 1,018 items comprise 395 side-talk recordings, 223 self-talk recordings, and 400 speaker-switch pairs, each mapped to one of three actions—silence, a tool call, or a natural-language answer. Existing evaluations usually pair well-formed requests with a response or tool label, so high tool-selection accuracy does not show that acoustic context controls execution; VGBench uses text-matched contrasts in which the same specified words license execution in one condition and silence in another, connecting acoustic-context use directly to action selection. Switch pairs hold the specified words fixed while changing trigger source, far-field rendering, and a 600 ms boundary; side-talk uses two consecutive utterances from the same stored speaker with balanced order, self-talk places command cores in planning, regret, quotation, sarcasm, and rhetorical-question frames, and all recordings passed two rounds of annotation, target-label verification, and script-audio consistency checks.

Off-the-shelf systems show a recognition-versus-gating gap: Step-Audio-R1.1 selects the target tool on 96% and 97% of same-speaker and text-only controls but mutes only 1% of switched commands, while Kimi-Audio-7B has the highest raw switch mute rate at 14% yet selects the target tool on only 53% and 44% of the two controls. These paired measurements locate the failure as not changing the action under a source-and-scene shift rather than a general inability to recognize tools; training-free adaptations change the error profile rather than solve gating—AURA lifts side-talk from 39.0% to 46.6% and self-talk from 57.0% to 63.2% but mutes no switched commands, and TwS raises self-talk muting to 85.7% while side-talk falls to 23.0%. Six raw audio LLMs and three training-free adaptations are evaluated under the same action contract, temperature zero, the same parser, and the same side-talk judge; the authors state these training-free implementations are local method-style adaptations rather than official full-pipeline reproductions, and Audio Flamingo 3 receives no canonical tool-selection credit because it emits its own bracketed action language.

In the VoxGate post-training case study, supervised fine-tuning reaches a 91.3% switch mute rate on the held-out set (73 of 80) while selecting the correct tool for every near-field same-speaker and text-only control; an exploratory GRPO stage reaches 92.5% switch muting (74 of 80), raises self-talk muting from 52.0% to 60.0%, and raises side-talk accuracy from 68.4% to 70.9%. This shows the benchmark's action boundary is learnable and is not bought by an always-mute policy, since near-field same-speaker and text-only controls expose false rejection; joint training also lifts the fixed 384-example WearVox protocol overall from 72.14% to 76.30%, above task-only SFT and GRPO at 63.28% and 67.71%. Post-training uses an approximately 80/20 item split within each scenario (320/80 speaker-switch pairs), only the training partition enters SFT or GRPO, and all reported scores come from the disjoint held-out partition; the authors state these are descriptive differences from one training run and that without an unpaired-GRPO or repeated-seed control the SFT-to-GRPO difference cannot be attributed to pair-aware grouping.

Factorized controls show the gate uses both source and distance: at fixed near-field rendering and a 600 ms gap, changing only the trigger source raises muting by 48.75 points for SFT and 52.50 for GRPO; at fixed source and gap, far-field rendering raises it by 58.75 points for both; the gap alone changes same-speaker near-field muting by only 1.25 points. The 91.3%–92.5% main switch result therefore combines source and distance dependence rather than isolated speaker identity, and the equal-budget cue-balanced SFT shows distance rejection is scene dependent—on the original validity WAVs it raises different/near muting from 35.00%/50.00% to 66.25%/82.50% and same/far muting from 60.00% to 90.00%, yet on scene-consistent eight-condition audio same/far muting remains only 3.75%/5.00%. The factorized study changes one cue at a time on the same 80 commands while holding text, tool target, and scoring fixed; cue-balanced SFT replaces only 1,212 speaker-switch audio slots in the 5,748-row mixture with other task proportions, 540 optimization steps, and the LoRA configuration unchanged, and the authors note the old and new conditions use different WAVs and should not be pooled as matched before/after measurements.

Perspective

This work targets evaluation and training for wearable and on-device voice agents: VGBench provides 1,018 action-level diagnostic items for comparing raw models, training-free adaptations, and post-trained models under one action contract, and it lets evaluators report false execution and false rejection separately. For designers, the findings point to an action gate that needs both source and proximity evidence, since "distance alone cannot reject a nearby bystander"; for trainers, supervised joint training learns the benchmark's conditional action mapping while preserving near-field and text-only tool controls and WearVox downstream tasks, making it a starting point for further gating work. The intended setting is the authors' conservative wearable authorization rule: a different source or a far-field trigger should be muted, while a same-source near-field trigger should call the tool.

Results rest on one split and one run per setting; the cue-balanced test varies wearer voice but not room, template, natural bystander speech, or open-set speakers; side-talk uses an LLM judge with no reported human-agreement estimate, which affects the free-response subset rather than the rule-scored mute and tool decisions; and VGBench tool accuracy checks tool names rather than arguments. The authors also note that cue-balanced SFT gains on the original validity WAVs and its behavior on scene-consistent audio come from different WAVs, so whether distance rejection is stable across acoustic scenes remains open, and without unpaired or repeated-seed controls the role of pair-aware grouping in the GRPO gain is not established. This summary is based on the paper's full text and its appendix tables and does not include reproduction experiments beyond the paper.

Sources