Skip to main content
Back to timeline
arXivSource publication:

ProbeGuard certifies abstention from how medical LLM consensus forms, answering 55.3% of the MedQA unanimous layer at 9.0% observed risk

Synopsis

The work introduces ProbeGuard, a certified abstention framework that decides from how multi-round consensus formed rather than from terminal agreement: it derives trajectory features from the system's execution log, checks unanimous first-round votes with rationale semantic entropy and an active counter-evidence probe, and converts these scores into a distribution-free selective-risk bound via stratified Learn-then-Test calibration; on MedQA terminal agreement scores 0.512 AUROC while the process score reaches 0.696, and 13.4% of unanimous votes are wrong, with the certified rule answering 55.3% of that layer at 9.0% observed risk.

Source-provided article image: Unanimously Wrong: Certified Abstention from How Medical LLM Consensus Forms
Figure 1 ·

Figure 1: Two unanimous consensuses are indistinguishable to snapshot-based abstention (a), yet one is knowledge-backed and one is inherited from a shared misconception. ProbeGuard (b) checks whether the rationales behind a unanimous vote support one another and stress-tests the consensus with retrieved counter-evidence, answering only where selective risk is certifiably bounded.

arXiv

Interpretation

Terminal agreement fails as an abstention signal in multi-round consensus systems: because the system stops deliberating once agreement is reached, the terminal vote is 1 for every converged question, so terminal agreement sits at chance on MedQA, MMLU-Pro medicine, and MedMCQA (0.512, 0.504, 0.507 AUROC). Prior agreement-based signals (self-consistency, semantic entropy, conformal abstention) are computed from a single round of samples; this work is the first to quantify this failure mode on a multi-round consensus system and to trace it to the early-stopping rule that makes the terminal vote constant. Across three benchmarks and three backbones (Qwen3-8B, HuatuoGPT-o1-8B, Llama-3.1-8B), terminal agreement is at chance in all seven cells, with the full grid in Appendix F.

The trajectory of consensus formation carries information absent from the terminal vote: the process score raises discrimination between correct and incorrect consensus from chance to 0.696 (MedQA), 0.704 (MMLU-Pro medicine), and 0.659 (MedMCQA), and yields the lowest selective risk on MedQA and MedMCQA. Trajectory features (trajectory-averaged agreement, majority-flip count, minority persistence, document-repeat rate) all come from logs the system already writes, at zero additional inference cost; ablations show only the rationale family is non-redundant, while the trajectory families are mutually interchangeable. 34 features across five families are fit by L2-regularized logistic regression under question-level cross-validation; the drop-one matrices in Appendix H show removing the rationale family lowers MedQA AUROC from 0.696 to 0.682, while removing any other family leaves it unchanged or marginally improved.

On the unanimous layer, where all candidates agree in the first round, every vote-based signal is constant yet the error rate is high: this layer covers 59% of MedQA with 13.4% wrong (100 of 746 votes), 17.7% on MMLU-Pro medicine, 22.3% on MedMCQA, and 79.0% on expert-frontier MedXpertQA. The work moves semantic entropy from answers down to the concluding claims of rationales and introduces an active counter-evidence probe: wrong consensus flips under counter-evidence at 14.0-26.5% while correct consensus flips at only 3.3-5.5%, an asymmetry stable across seven cells. Control arms attribute the flips to the retrieved content itself: the devil's-advocate instruction alone flips none of the 746 votes, irrelevant documents flip 1.0% of wrong consensus, and counter-evidence under a neutral instruction retains nearly the full asymmetry (12.0% versus 3.6%).

Stratified Learn-then-Test calibration turns the process scores into a distribution-free selective-risk bound: on the MedQA unanimous layer the combined score certifies answering 55.3% of questions at 9.0% observed risk over 50 splits, about 32% of all MedQA questions; augmenting the calibration side with in-domain unanimous questions lifts certified coverage to 89.2% and certifies 21.9% at a stricter target for the first time. Prior medical abstention methods provide heuristic scores whose thresholds carry no error guarantee; certifying per process stratum yields non-zero certified coverage even on the stratum where vote-based signals are constant. The per-split violation fraction is 0 of 50 for every Learn-then-Test arm reported, meaning no split's realized test risk exceeds its target; without stratification, naive quantile calibration misses its target on MedQA (15.7% realized risk) and global Learn-then-Test returns no certified threshold at any level swept.

Perspective

The framework targets question-answering systems that deliberate in rounds and write an execution log; a deploying hospital can calibrate ProbeGuard on its own case mix and refer every question outside the certified region to a clinician, with each referred question carrying its challenge documents and the record of dissent so the clinician sees why the system withheld an answer. The guarantee applies in any stratum whose base error rate is at or below the target risk; a stratum above it receives no certified region. Tighter certificates require more calibration data, and in-domain unanimous questions accumulate as the system runs. Appendix I shows the same calibration can be re-run within specialty strata, with emergency medicine erring about twice as often as the layer average and warranting more conservative routing.

Evaluation is limited to multiple-choice benchmarks, so whether the flip asymmetry and rationale signals carry over to free-text clinical questions is untested; the three backbones are 8B-scale models on a single published substrate, and the signal definitions transfer but have not been measured elsewhere. The certificates hold under an i.i.d. within-stratum assumption, so a deployment must calibrate on its own case mix and recalibrate when that mix drifts. Whether clinicians find the returned dissent records useful has not been evaluated; a prospective evaluation with clinicians is the next step. In addition, several equations and the hyperparameter table are not fully rendered in this parse, so readers needing exact values for reproduction should consult the original appendices.

Sources