Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight
Synopsis
This study compares human judgments with GPT-4.1 and GPT-5 as LLM judges on telecom and retail voice-agent conversations under three configurations, p0 (no persona), p1 (static persona), and p2 (dynamically inferred context), and finds that LLM judging holds stable metric-level trends and often follows human relative assessments while absolute scores diverge, with the largest gaps on safety metrics (IAS, SR) and Recovery Turn Count, supporting LLM judges as a scalable first-pass component within a hybrid pipeline that keeps human oversight.
Fig. 1: Workflow for comparing human and LLM-as-Judge evaluations, from voice-agent conversation scoring through divergence analysis, calibration, and hybrid evaluation design.
arXivInterpretation
Across 242 telecom and retail voice-agent conversations, LLM judges and human raters keep stable relative metric trends, but their absolute scores differ systematically. Prior LLM-as-a-Judge work focused largely on text dialogue, position bias, and prompt sensitivity; this paper scores the same voice-agent interactions under p0, p1, and p2 with GPT-4.1, GPT-5, and three human annotators, and reports human-to-judge ratio tables. Evidence comes from the human-to-LLM ratios in Tables II and III: TE stays near 1 across both judges and all three configurations, CR stays near 1.0 to 1.1, while IAS and SR exceed 3 in Telecom; the sample is 242 conversations across six configurations.
Safety metrics show the largest divergence: humans record far more safety events than either LLM judge, and in Telecom this gap does not depend on the specific annotator. The paper links the safety divergence to the companion benchmark MM-τ-p2, where both GPT-4.1 and GPT-5 produced low and inconsistent safety precision and recall, and notes that necessary escalation (such as SIM-lock cases) is hard to state unambiguously in a rubric, so judges swing between two readings of the same behavior. In Telecom all three annotators give IAS from 0.794 to 1.000 and SR from 0.779 to 1.043, well above the judge level of roughly 0.15 to 0.25 implied by the Table II ratios; Retail safety ratios are smaller (IAS 1.189 to 1.754, SR 1.056 to 1.626), and in at least one configuration the direction depends on the reference annotator.
Recovery Turn Count (RTC) is among the largest divergences, with LLM judges often producing values below one while every human annotator reports above one in every domain and configuration. The paper frames RTC as a multi-step tracking task that requires tracing an error back to its origin and counting forward to resolution, and cites expert-LLM agreement ranging from a weighted kappa of 0.17 to 0.86 across 21 dimensions in empathic communication work, plus industry reports of agreement below half on open-ended multi-step reasoning. In Telecom the two judges implied by Eval 1 and the ratios fall between 0.22 and 1.22 turns, below every annotator value in the matching cell; in Retail, however, the implied GPT-5 estimate exceeds at least one annotator in each configuration, with p0 at about 1.75 turns against 1.385 for Eval 1.
Human-judge divergence and annotator disagreement fall on different metrics, so the feasibility of calibration varies by metric. The paper uses mean relative range to separate two kinds of uncertainty: ARGA (0.77), RTC (0.46), and CP (0.27) are the least reproducible across annotators, while CFA (0.11), TE (0.14), CR (0.15), and UES (0.15) are the most reproducible; the widest human-judge gaps are instead on IAS and SR. Based on the range across three annotators divided by their mean, averaged over six domain-and-configuration cells; the paper also reports pairwise Pearson r=0.975 between Eval 1 and Eval 3 and r=0.903 between Eval 1 and Eval 2 (N=54), noting that Retail Eval 3 is identical across configurations and inflates that correlation.
Perspective
The framework targets conversation-level quality and safety evaluation of telecom and retail customer-service voice agents, and suits teams that want LLM judges to carry scalable first-pass assessment while humans handle safety and recovery judgments. The paper reports that calibration is well supported in Retail (Pearson r above 0.9 for both judges) but weaker in Telecom, where it is not significant for GPT-4.1, so setting calibration and judge choice per metric and per domain is the mode of use this result supports. The paper also proposes extending the framework to audio signals, speaker-turn boundaries, and timestamp data to cover response windows, interruptions, overtalk, and barge-in.
Agreement is measured on metric-level aggregates, which the paper describes as an upper bound on conversation-level agreement, so a conversation-level reliability coefficient remains an open observation. Retail Eval 3 is identical across p0, p1, and p2, so cross-configuration comparison in that domain rests on Eval 1 and Eval 2. The paper did not recompute ratios or correlations under Eval 2 or Eval 3, so the size of the human-judge gap on RTC, ARGA, and CP moves with the reference annotator. Rubric wording for necessary escalation and multi-turn recovery remains ambiguous, and tightening those rubrics is a direction the paper proposes. The current benchmark also does not model voice-specific phenomena such as response windows, prolonged silence, interruptions, overtalk, and barge-in, which require acoustic and timing information.
