Skip to main content
Back to timeline
arXivSource publication:

BayesJudge models conflicting human and LLM judgments as posterior distributions, raising expert prediction accuracy to 65.3% on SummEval

Synopsis

The authors propose BayesJudge, an online Bayesian meta-evaluation layer for conflicting human-LLM judgment streams that outputs a posterior verdict distribution per comparison while estimating rater-specific human confusion matrices and LLM presentation-order bias, using tie-open labels and paired AB/BA order swaps as protocol-level identification instruments; synthetic experiments recover prespecified evaluator parameters, and on SummEval it detects LLM position effects, separates expert from crowdworker behavior signatures without rater metadata, and yields posterior uncertainty that correlates with human disagreement.

Source-provided article image: BayesJudge: Uncertainty-Aware Bayesian Meta-Evaluation of Human and LLM Judgments
Figure 1 ·

Figure 1 : Motivating example. BayesJudge transforms heterogeneous, biased pairwise evidence into a probabilistic posterior verdict and diagnostics, rather than a single hard label. (a) Experts (E), crowdworkers (C), and LLM judges ( J A ​ B J_{AB} , J B ​ A J_{BA} ) provide conflicting observations; the simplex shows each channel’s distribution over true verdicts. (b) Hard reductions (majority vote, single-order LLM, swap heuristic) discard ambiguity and bias. (c) BayesJudge centers on latent π t \pi_{t} , jointly modeling human confusion C r C_{r} and LLM position bias β j \beta_{j} for streaming inference. (d) Outputs: posterior verdict, rater reliability C r C_{r} , LLM bias β j \beta_{j} , and entropy-based uncertainty H H .

arXiv

Interpretation

BayesJudge treats conflicting human and LLM judgments as evidence rather than noise, producing a posterior verdict distribution per comparison (with a tie or ambiguity state) while jointly inferring rater-specific confusion matrices and LLM presentation-order bias. Prior LLM-judge debiasing and human-annotator noise modeling formed two disjoint lines that both output single-point aggregate estimates; this work models both bias sources jointly inside one online framework and retains uncertainty instead of collapsing to a hard label. The paper formulates an exact online posterior recursion and approximates it with a Rao–Blackwellized assumed-density SMC filter; main experiments use the hard-MAP Dirichlet update, with the soft update reported as an ablation.

Tie-open labels and paired AB/BA order swaps act as protocol-level identification instruments: the former keeps item-level ambiguity observable, the latter separates item preference from presentation bias. The paper proves as propositions that forced-choice-only labels cannot identify the tie parameter, and that with a single order item preference and position bias are confounded; this elevates order swapping from a robustness heuristic to a provable identification design. On the synthetic stream, tie-open labels drive tie-parameter MSE down and bias toward zero, whereas forced choice leaves MSE an order of magnitude higher with non-converging bias; the single-order likelihood yields a markedly biased mean estimate while the swap likelihood gives substantially tighter posterior variance convergence.

On SummEval the joint model shows a pronounced rater-type reversal: the human-only channel predicts crowdworkers well (68.9% Acc) but fails on experts (43.7%), the LLM swap channel better captures expert judgment (59.5%), and fusion achieves the best expert prediction (65.3% Acc) and the lowest crowdworker calibration error (ECE 0.042). MACE matches the human-only channel on crowdworkers (69.10% Acc) but collapses on experts (38.00% Acc), while Bradley–Terry–Davidson and GLAD sit in between on both groups; this work keeps an independent parameter space per evaluator so heterogeneous biases mutually constrain one another during fusion. Evaluated under a leave-one-human-out protocol with 8 raters (3 experts, 5 crowdworkers) and GPT-4o-mini dual-order judgments, reporting Acc, NLL, Brier and ECE.

Posterior entropy correlates positively with genuine human disagreement and can serve as a routing signal. Majority-style aggregation reports only a winner and no item-level uncertainty; this work makes posterior entropy the default uncertainty score and notes it can be high when humans appear unanimous but LLM swaps reveal hidden ambiguity. The lowest entropy quintile (Q1) shows minimal human disagreement and the highest quintile (Q5) near-maximal disagreement, with an especially pronounced Q1-to-Q2 jump; middle buckets are not strictly monotone, which the paper attributes to conditioning on LLM swap evidence and inferred rater reliability.

Perspective

The framework targets streaming pairwise evaluation settings where tie-open labels can be collected and AB/BA dual-order calls can be issued, such as crowdsourced leaderboards and automated benchmarks. For researchers and engineering teams wanting item-level uncertainty, rater failure modes, and judge position-bias audit signals, it provides directly usable diagnostic outputs; the paper's protocol checklist maps tie labels, order swaps, repeated human ratings and repeated judge calls to the quantities each identifies. Synthetic experiments verify parameter recovery and structural correctness, while SummEval experiments verify diagnostic validity and predictive reliability, giving the two a clear division of labor.

In real evaluation data the limited number of human annotations per sample, compounded by annotators not being consistent across samples, constrains precise modeling of evaluator bias patterns; the work currently considers only LLM position bias, while factors such as response length also influence judgments in practice; and the downstream utility of the posterior verdict distributions remains to be directly validated. In addition, the theoretical guarantees refer to the exact posterior target, whereas the implementation uses a MAP-profile potential and hard-MAP Dirichlet update as scalable assumed-density approximations, a distinction readers should note when exact coverage is required.

Sources