Skip to main content

Daily report

AI and science frontiers · 2026-08-25

Only content delivered through the publication boundary on this date is included.

arXiv

Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight

This study compares human judgments with GPT-4.1 and GPT-5 as LLM judges on telecom and retail voice-agent conversations under three configurations, p0 (no persona), p1 (static persona), and p2 (dynamically inferred context), and finds that LLM judging holds stable metric-level trends and often follows human relative assessments while absolute scores diverge, with the largest gaps on safety metrics (IAS, SR) and Recovery Turn Count, supporting LLM judges as a scalable first-pass component within a hybrid pipeline that keeps human oversight.