Skip to main content
Back to timeline
arXivSource publication:

When a Judge Learns to Say “I’m Not Sure”: Decision-Only Judging and a Confidence Cascade

Synopsis

The study compares a decision-only judge, JEV, which returns a verdict plus label probabilities, against sixteen generative and reward-model judges on preference, factuality, and answer-adjudication tasks with blinded human adjudication, finding it within three percentage points of the strongest comparator, GPT-6, on ordinary preference and evidence-grounded factuality at about 0.36% of that comparator's fee, with the gap concentrated in low-confidence decisions, so a frozen cascade that accepts confident verdicts and escalates uncertain ones retains about 99% of the comparator's accuracy at lower cost.

AI-generated editorial illustration: JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

Interpretation

A decision-only judge approaches the strongest generative judge on ordinary preference and evidence-grounded factuality while costing one to two orders of magnitude less. Prior efficiency discussion around LLM-as-a-judge centered on generative judges and routing algorithms; this work supplies a measured operating profile of a hosted decision-only interface: JEV scores 92.2% on RewardBench versus GPT-6's 93.5% and 87.5% on HaluEval versus 86.7%, at a median latency of 0.152 seconds and about $0.044 per 1,000 judgments, against 1.885 seconds and $12.182 for GPT-6. 1,312 base judgments across seventeen configurations with source-question cluster bootstrap intervals, plus blinded human adjudication of every item where the two judges' correctness differed; the authors note the adjudication was completed by one team member on all items rather than being consensus at scale.

The gap is not uniform: it concentrates where a judgment requires checking a multi-step derivation or resisting a more elaborately written wrong answer. This turns “is a cheap judge good enough” into a workload profile: on JudgeBench JEV scores 78.6% against GPT-6's 93.1%, with reasoning at 68.4% versus 95.9% and coding at 76.2% versus 97.6%; on RM-Bench it scores 84.0% on matched-style pairs but 74.8% when the rejected answer is the more elaborate one. 350 JudgeBench pairs and 480 judgments per RM-Bench condition, all paired comparisons with source-cluster intervals; human adjudication sided with GPT-6 on 57 of 69 disputed items and with JEV on one.

The judge's confidence orders its own errors, which supports automatic escalation. Rather than proposing a new routing algorithm, the paper gives a frozen threshold policy: judge each pair in both orders, average the aligned probability, accept when confident and hand uncertain items to a stronger model. Across 990 base-order judgments, accuracy is 47.7% below maximum label probability 0.5, 76.5% in 0.5–0.8, 93.9% in 0.8–0.95, and 99.1% at 0.95 and above; the frozen two-order policy accepts 53.7% of 510 extension pairs, scores 92.5% against GPT-6's 93.1%, and uses 56.8% of the fallback's fee.

The confidence signal works only where the first-stage judge is competent but uncertain, and fails where it is confidently misled. This bounds the cascade's applicability: on RM-Bench hard pairs the AUROC of confidence against correctness falls to 0.770 and the cascade retains only 96.5% of GPT-6's accuracy; on reference-free prose the AUROC is 0.518 and no threshold helps. Paired comparisons on the same prompts across nine style combinations in both orders; the authors state these thresholds are read off the items they are scored on, and only the frozen two-order policies test a pre-specified rule.

Perspective

The result applies to evaluation workloads that come with a reference or an evidence passage, or that are ordinary preference judgments: there a decision-only judge approaches the strongest judge at far lower fee, and confidence can route uncertain items to a stronger model. It does not apply where a verdict requires checking a multi-step derivation, resisting a more elaborately written wrong answer, or judging natural prose without a reference, where the paper recommends escalation or local validation first. The authors' deployment recipe is to judge pairs in both orders and average the aligned probability, choose the escalation threshold on a local selection set and re-check it on held-out items, count invalid outputs as errors, treat confidence as an escalation signal rather than a certificate, and run a small local validation before extending to a new workload.

Several open questions remain for a careful reader: the multi-family comparison followed inspection of the original study and is exploratory; the human adjudication covers 183 items selected by judge disagreement rather than a random audit, was completed by one team member on all items with an author adjudicating disagreements, and 36 final labels are indecisive; temperature scaling did not transfer across workloads and fitted temperatures point in different directions; the single-order cascade thresholds are read off the items they are scored on, and only the frozen two-order policies test a pre-specified rule; latency comes from one client location and collection window and fees are estimates rather than invoices; training overlap and benchmark contamination are unknown, and specialized professional domains are untested. The loaded text is the full paper, but some interval values in the tables are not fully rendered in the text, so exact intervals should be checked against the original tables.

Sources