Skip to main content
Back to timeline
arXivSource publication:

VoxParity tests 28 voice agents on 183 scenarios: when audio calls for protection, every system more often executes the routine request (41% against 12%)

Synopsis

The work introduces VoxParity, a benchmark of 183 scenarios from 14 sectors in which one transcript stays fixed while the audio changes (a coaching voice, a medical monitor beeping, a mayday under a radio check, noise over a drug name, a child's voice placing a bet, a frightened whisper) and the correct executable tool call changes with it; across 206 cue-bearing cells, all 28 systems (including nine production realtime agents) execute the routine request on protective calls at 41% against 12% over-triggering on clean calls, and only 11 of the 23 systems with a transcript path pass the words-only null test, with passes coming almost entirely from items that state the rule.

AI-generated editorial illustration: Almost Human, Except When It Matters: VoxParity and the Decisions a Voice Should Change

Interpretation

The work builds a benchmark scored on the typed tool calls agents execute: 183 scenarios from 14 sectors, seven kinds of audible cue plus a channel control, 28 systems (13 file-mode API models, 9 production realtime agents, 6 local open-weights models), a judge-free deterministic scorer, and about 20 volunteer players choosing among the same actions on the same audio as a reference. Prior studies observed that realtime agents act on the words rather than the voice in 3 scenarios at 5 trials per cell, or found that adding audio to the transcript barely changed open 7B audio models' decisions; VoxParity turns these observations into a scaled measurement on executed calls against a words-only null. The gold flips in 130 of the 132 items with two runnable variants; menus are shuffled per item with a fixed seed so that a policy of picking the second tool whenever the caller sounds off, which would score 0.977 on unshuffled menus, gains nothing; scoring follows the Berkeley Function Calling Leaderboard's syntax-tree matching, requiring the tool name and each typed argument to match.

The direction of errors runs toward the words: all 28 systems, from 11 vendors and three serving modes, execute the routine action on protective calls more often than they over-trigger on clean ones, 41% against 12% pooled, and 34% against 18% for the four leading systems, while the words-only cascade carries out the routine request on 58%. The work separates hearing the audio from acting on it: against each system's own transcript, hearing the audio lowers unsafe execution by 0.12 but leaves over-triggering unchanged (+0.00), and the direction never reverses. A difference-in-differences on identical cells clustered by item across the 23 systems with a transcript path, with the drop significant for 11 after Holm correction; the direction holds with any one vendor removed; an audit of 60 sampled classifications agreed with each item's gold rationale in 57.

Acting on the caller's state happens mainly where the item states its rule: the four leading systems gain +0.33 over the null on emotional cells where the rule is stated and show no detectable gain (+0.00) where it is not; among heard feelings, they take the words' action on 6% of acute alarm (panic, gasping, grief, whispered duress) against 52% of quiet states (confusion, resignation, slurring, tears, sarcasm). The work separates cues that change the facts a request depends on from cues that change what the caller's state asks of the agent, and reports each contrast by whether the rule is stated, locating where the gap sits. These are exploratory analyses that the authors state were chosen after looking at the data; across the 23 systems, 18 clear zero on stated-rule items against 1 on items that leave the rule implicit (unadjusted, observational).

At the frontier the bottleneck is the bridge from hearing to deciding: perfect hearing would add 0.04 credit and perfect deciding 0.28; describing the caller's delivery in the prompt (an upper bound built from the variant's specification) and stating the rule each recover part of the gap, and with both gemini-3.7-flash acts correctly on 0.83 of emotional calls whose rule was unstated, still 0.17 below other cues. The work carries Hear2Act's finding that an audio-inferred state written into text lifts action onto typed tool calls, and shows the description must carry an intensity that a model's own description understates. Both notes are upper bounds built from the rendering specification rather than the clip; the model's own pre-action descriptions often understate the state it heard (30 of about 50 emotional calls understated, denied or mislabelled it, and on 22 of the 30 its own forced-choice probe had identified the cue).

Perspective

The work addresses deployers, auditors and buyers: the words-only null test can be run on the development split (40 items, 81 cells) to screen for large effects and at full scale for a verdict, by running each cell as audio and as transcript, running a words-only pipeline on the same cells, and passing the agent only when its paired audio-minus-transcript change clears the pipeline's. It applies to voice-agent settings where sector rules require acting on audible cues, such as emergency calls, fraud prevention, aviation and maritime radio, gambling age verification and vulnerable-customer protection. The authors also suggest that builders state their rules and evaluate, and train, tool-calling policies on pairs in which the same words call for different actions.

A careful reader would still watch that only the null-test verdicts are confirmatory, while the error direction, the facts-versus-feelings pattern, the hearing-versus-deciding account and the comparison with other measures were chosen after looking at the data; that the emotion results rest almost entirely on one Gemini-TTS voice (298 of 309 scored cells use a single voice); and that the human reference comes from about 20 self-selected unpaid volunteers, one of whom supplied 23% of the answers, who answered the perception question after locking their action. Stimuli are synthetic, and all human recordings come from the author, who knew the study's hypothesis. The protocol mapping is the authors' reading of written rules rather than a legal finding, and 19 legacy items lack written grounding. The roster covers the vendors and realtime APIs available in September 2026 except Amazon's Nova 2 Sonic, and not every model size or generation. The loaded text is the full paper, but some tables and appendix details arrive as text, so specific numbers are best checked against the original and the released repository.

Sources