Freezing debate prevented 29 collapses but lost 108 corrections across 6,925 MMLU-Pro debates
Synopsis
The work introduces an auditable protocol for homogeneous three-agent LLM debate that decomposes final accuracy into preserved, collapse, correction, and unrepaired transitions with collapse onset and signed intervention utility, identifies 253 collapses across 6,925 MMLU-Pro debates, and shows via replay that a leave-one-model-out probe-gated freeze prevents 29 collapses while losing 108 corrections under equal weights, so collapse prevention alone can recommend the wrong policy.
Interpretation
The work replaces "did the final answer improve" with a directional transition ledger: collapse means an initially correct majority ends wrong, correction means an initially wrong majority ends correct, and each run also records collapse onset and signed intervention utility. Standard final-accuracy evaluation conflates two opposing mechanisms; the authors note that Sonnet 4.5 and Llama-3.1-8B have nearly identical debate flip rates, yet Llama's conditional collapse rate is about four times higher, so the missing variable is direction, not motion. 253 collapses are identified across 6,925 MMLU-Pro debates alongside a parallel correction ledger; all labels are rule-based with no LLM judge or human coder, and parser-failure and tie-rule sensitivities are audited.
Replay experiments expose the central tradeoff: a leave-one-model-out probe-gated freeze prevents 29 collapses but loses 108 corrections under equal weights, for negative net utility, and simple Round-1 majority-change gates and learned Round-1 stumps are also negative. This changes how interventions should be judged: looking only at collapse prevention would recommend the wrong policy, so the authors argue any action policy must be scored by held-out signed replay under declared weights. A frozen leave-one-model-out aggregate over debates for the probe-gated policy, plus Round-1 replay rows on a separate trace cohort; at a collapse weight of 10 the probe-gated freeze is still net-negative while the Round-1 majority-change rule turns positive.
A compact pre-debate 8-probe screen works as a triage signal: its unadjusted family-level association with conditional-collapse risk is high (G=7, Spearman rho=0.893, exact two-sided p=0.0123), but initial-majority accuracy is a close free comparator (rho=0.821; family partial rho=0.767, p=0.0877). The authors explicitly do not treat it as calibrated or capability-adjusted prediction; it is meant to prioritize which model-scaffold rows deserve expensive trace logging, and once initial answers are observed, disagreement is the stronger runtime diagnostic. The probe measurement is itself repeatable, with reported split-half reliability, test-retest reliability, and inter-agent ICC; model-row Spearman, leave-one-family, taxonomy, and errors-in-variables bootstrap checks bound rather than enlarge the claim.
Round-level traces localize most collapses to the first debate round: 149 of 253 collapses begin in Round 1, and adding Round-1 trajectory features raises pooled out-of-fold AUC from 0.573 to 0.658, rising to 0.673 with Rounds 2 and 3. This is runtime localization rather than causal mediation: early disagreement can precede both harmful cascades and useful recovery, so Round 1 is a diagnostic substrate, not a standalone deployable gate. Paired bootstrap confidence intervals over debate indices, plus a four-round trace-schema replication; in a Bayesian multilevel logistic model the model-level slope has a highest-density interval crossing zero, which the authors read as trajectory localization rather than causal mediation.
Perspective
The protocol targets a deliberately narrow setting: homogeneous, closed-book, three-agent MCQ debate, with MMLU-Pro as the primary benchmark. Its reusable output is a row schema, rule-based coder, parser and stability audits, cost accounting, release flags, and zero-API rebuild scripts, so a new model-scaffold row can be compared under the same denominators and signed utility ledger. For a reader, this means that when evaluating a debate, judge-debater, abstention, or deferral system, one can first build the initial-majority to final-majority transition table, report conditional collapse, correction, and onset, and then score any gate by held-out signed replay under declared collapse/correction weights; the probe screen fits decisions about which model-scaffold rows deserve expensive trace logging, while Round-1 trajectory features fit runtime diagnosis once debate has begun. The authors also tier the release: aggregate tables and code are open, while convince-wrong probe templates and full transcripts are gated for research use, so reusers must state explicitly which matrices are missing or gated.
Several open questions remain for a careful reader. The screen's marginal value is limited: initial-majority accuracy is a close free comparator, and the family-level partial after controlling for initial accuracy is no longer confirmatory, so it is better suited to triage than to per-question prediction. Round-1 localization is diagnostic rather than causal mediation, early disagreement precedes both harmful cascades and useful recovery, and converting the signal into net accuracy gains still requires model-aware cost functions and complementary strategies. Outside the scope are heterogeneous judge-debater systems, tool use, retrieval, free-form generation, rubric- or judge-scored answers, and agent workflows that gather evidence over longer contexts; mixed-model panels and GSM8K rows are feasibility checks only. Reasoning-mode rows show collapses and corrections both persist, but do not establish that debate outperforms a compute-matched solo reasoning baseline. In addition, the per-debate Round-1 feature matrix and the strict matched DisagreementGate matrix remain gated, so those results are diagnostic, closed-API rows may drift as providers update models, and robustness of the screen to the probe reply temperature remains unresolved.
