Only content delivered through the publication boundary on this date is included.
arXiv This study compares human judgments with GPT-4.1 and GPT-5 as LLM judges on telecom and retail voice-agent conversations under three configurations, p0 (no persona), p1 (static persona), and p2 (dynamically inferred context), and finds that LLM judging holds stable metric-level trends and often follows human relative assessments while absolute scores diverge, with the largest gaps on safety metrics (IAS, SR) and Recovery Turn Count, supporting LLM judges as a scalable first-pass component within a hybrid pipeline that keeps human oversight.
This study compares human judgments with GPT-4.1 and GPT-5 as LLM judges on telecom and retail voice-agent conversations under three configurations, p0 (no persona), p1 (static persona), and p2 (dynamically inferred context), and finds that LLM judging holds stable metric-level trends and often follows human relative assessments while absolute scores diverge, with the largest gaps on safety metrics (IAS, SR) and Recovery Turn Count, supporting LLM judges as a scalable first-pass component within a hybrid pipeline that keeps human oversight.
This study compares human judgments with GPT-4.1 and GPT-5 as LLM judges on telecom and retail voice-agent conversations under three configurations, p0 (no persona), p1 (static persona), and p2 (dynamically inferred context), and finds that LLM judging holds stable metric-level trends and often follows human relative assessments while absolute scores diverge, with the largest gaps on safety metrics (IAS, SR) and Recovery Turn Count, supporting LLM judges as a scalable first-pass component within a hybrid pipeline that keeps human oversight.
This study compares human judgments with GPT-4.1 and GPT-5 as LLM judges on telecom and retail voice-agent conversations under three configurations, p0 (no persona), p1 (static persona), and p2 (dynamically inferred context), and finds that LLM judging holds stable metric-level trends and often follows human relative assessments while absolute scores diverge, with the largest gaps on safety metrics (IAS, SR) and Recovery Turn Count, supporting LLM judges as a scalable first-pass component within a hybrid pipeline that keeps human oversight.