Pre-verdict residual stream activations predict LLM judge position flips with AUROC .621–.850 and transfer to .685–.853
Synopsis
Using nested grouped cross-validation on 534 JudgeBench pairs and 1,802 MT-Bench comparisons, this work trains L2-regularized linear probes on residual stream activations recorded immediately before the verdict to predict position flips, where swapping candidate order changes the judge's selected response identity; AUROCs reach .621–.850 across four judges, exceeding a combined baseline of verbalized confidence, verdict logits, length differences, and the initial choice by .062–.113, and frozen probes transfer to MT-Bench at .685–.853.
Interpretation
Pre-verdict residual stream activations predict position flips: AUROCs range from .621 for Llama-3.1-8B to .850 for Qwen3-8B across 534 shared JudgeBench pairs, with all confidence intervals excluding .5. Identifying a position flip conventionally requires judging each pair in both orders, which detects positional sensitivity only after the judgment; this work instead reads an activation during a single evaluation, immediately before the verdict token, to predict which comparisons are susceptible. Nested grouped cross-validation with five outer and three inner grouped folds keeps variants of the same question together, and feature standardization and model selection occur only within training folds; each comparison yields exactly one held-out prediction; 95% confidence intervals use 2,000 group-bootstrap resamples with the underlying question as the resampling unit.
The linear probe outperforms a combined baseline using verbalized confidence, verdict-label logits, response-length differences, and the judge's initial choice, with AUROC gains of .062–.113 and every paired confidence interval excluding zero. Because the baseline already includes the judge's choice, the improvement is not explained by choice alone; under the stricter choice-matched AUROC, which compares flipping and non-flipping examples with the same choice, the probe still improves by .108–.126 with all paired intervals excluding zero. Baselines are evaluated on the same held-out folds as the probe and follow the same nested fitting procedure; the authors also report paired AUPRC and Brier differences, noting that the Qwen3-4B AUPRC interval and the Llama-3.1-8B Brier interval include zero.
Linear probes trained on JudgeBench and then frozen generalize to MT-Bench without adaptation, achieving AUROCs of .685–.853 with every 95% confidence interval above .5. No fitting, threshold selection, or recalibration uses MT-Bench data, indicating the predictive signal is not confined to a single response-pair distribution. Hyperparameters are selected by grouped cross-validation over all 534 JudgeBench pairs, after which one linear probe and its standardization are refit and frozen on the complete JudgeBench set; MT-Bench PFR ranges from .172 to .379, AUPRC exceeds flip prevalence for every judge, and Brier score is lower than a constant predictor fixed to JudgeBench prevalence.
A repeated-presentation control finds no same-order verdict variation: on a predeclared 100-pair subset, repeating the same order changes no verdict, whereas swapping order changes 19–47 verdicts depending on the judge. This control separates observed flips from same-order noise under deterministic inference, supporting candidate order rather than sampling randomness as the source of the flips. The subset is selected before inference using a seeded hash of the pair identifier; all judgments are generated deterministically in bfloat16 with greedy decoding and a 256-token generation allowance.
Perspective
The result applies to open-weight judges whose residual streams are accessible: four relatively small models (Qwen3-1.7B, Qwen3-4B, Qwen3-8B, and Llama-3.1-8B-Instruct) on two response-pair distributions. It enables scoring a pair's flip risk from a single evaluation, which could concentrate two-order judging on high-risk pairs; the authors note that untested settings still require evaluating sampled pairs in both orders. Probe scores predict position-flip risk rather than general reliability.
Linear decodability does not establish that the model uses the decoded information causally or explicitly estimates positional susceptibility, and residual stream access excludes closed models. The shared analysis sets retain 86.1% of the JudgeBench pool and 75.1% of the MT-Bench comparisons, and inclusion requires valid outputs in both candidate orders for every judge, so results are conditional on format compliance. The combined baseline does not exhaust all measures of difficulty or uncertainty, the confidence-before-verdict format may weaken verbalized confidence, and outperforming the fitted linear baseline does not exclude nonlinear predictive relationships among its features. The authors do not evaluate a risk-threshold policy, realized judgment savings, or net runtime including activation extraction.
