SAUCE estimates uncertainty in parallel multi-agent reasoning via filtering-style sequential inference, improving misclassification detection, selective prediction, and calibration across five backbones and two protocols
Synopsis
The work proposes SAUCE, a training-free uncertainty estimator that casts uncertainty in parallel multi-agent reasoning as sequential inference over a latent system-level belief: each round's cross-agent agreement serves as evidence and its generation uncertainty sets the evidence weight, updated through a filtering-style recursion, improving misclassification detection, selective prediction, and calibration across five backbones, five benchmarks, and the Debate and DyLAN protocols.
Figure 1 : Overview. (A) Two MAS trajectories can reach the same final answer but differ greatly in reliability: a stable run whose outputs converge across rounds (top) is more trustworthy than an unstable run whose outputs oscillate (bottom). A static view that only considers the final output cannot distinguish them, motivating uncertainty estimation from interaction dynamics. (B) SAUCE models multi-round interaction as a sequence of round-level observations ( M t , R t ) (M_{t},R_{t}) , where agreement is weighted by generation uncertainty, and performs a filtering-style update over a latent system-level belief. The resulting trajectory summary defines the SAUCE uncertainty score u ( 𝒙 , τ ) u({\bm{x}},\tau) (Eqn. 11 ).
arXivInterpretation
The paper reformulates uncertainty in parallel multi-agent reasoning systems as sequential inference over the interaction trajectory rather than a static score on the final output. Prior single-LLM uncertainty methods (sampling-based, Bayesian, internal-signal) do not directly address multi-agent structure; the paper argues a static summary is blind to the temporal order of evidence, so two trajectories with identical average signals but opposite dynamics receive the same score. The argument is supported by a formalization in which the trajectory is a sampled multi-round interaction and uncertainty is a state updated recursively per round, with a probabilistic graphical-model interpretation in Sec. 4.
SAUCE is driven by two round-level signals available at no extra cost: modal-answer share as cross-agent agreement, and mean token-level predictive entropy across agents as generation uncertainty. Agreement supplies the evidence while generation uncertainty determines how strongly that evidence updates the belief, merging cross-agent disagreement and per-agent generation uncertainty into one system-level estimate. Both signals are explicitly defined (modal-answer share; token entropy from top-k log-probabilities), and the paper reports that correct samples consistently exhibit higher average confidence than incorrect ones.
The recursion is shown to be the exact posterior update of a linear-Gaussian state-space model, with the round-level estimate as the posterior mean of the latent system belief and the stability weight as its posterior variance. This gives the filtering-style update a probabilistic reading rather than a heuristic weighting: agreement acts as a noisy observation of the latent belief and generation uncertainty as its observation-noise variance. Proposition 4.1 gives closed-form posterior mean and variance, proved by induction in Appendix D.1, and Appendix D.3 shows the objective is simultaneously the exact marginal likelihood and a tight ELBO.
Across five backbones, five benchmarks, and two MAS protocols, SAUCE leads most cells and every per-backbone average, and stays more robust than last-round log-likelihood as debates lengthen. Single-LLM scores (PE, LL) frequently land near or below chance under MAS, and specialized MAS estimators (UDPO, MATU) lead only on isolated (model, protocol) cells; ablations indicate the gain comes from the sequential update rather than a static combination of agreement and generation uncertainty. The main tables report a full AUROC and AUARC grid, calibration uses Brier after Platt scaling with 5-fold cross-validation, and prompt-level bootstrap plus repeated-rollout means and standard deviations are provided; on GPT-5.4-mini / BBEH-mini SAUCE is near chance.
Perspective
The result targets parallel multi-agent reasoning systems, where agents give comparable answers to the same problem each round and a final prediction is produced by an aggregation function such as majority vote; Debate and DyLAN are representative protocols. In this setting SAUCE needs only parsed answers and top-k token log-probabilities already present in a standard rollout, requiring no extra sampling, model weights, or hidden states, so it carries over to closed-source models that expose only top-k log-probabilities. Its hyperparameters are scalars tuned on a single calibration cell per MAS protocol and transferred across models and datasets. For downstream users, the Platt-scaled per-trajectory confidence can serve as a gating signal for selective prediction, abstention, or uncertainty-aware orchestration, and retaining earlier-round evidence is more stable than last-round confidence as debates lengthen.
On GPT-5.4-mini / BBEH-mini, SAUCE's AUROC is near chance and the repeated-rollout mean is likewise near chance, so its relative advantage there should not be read as reliable discrimination. The association between interaction history and correctness in Appendix C.6 is descriptive; the paper states it does not establish a causal effect of switching and is not a calibrated correctness probability. The hyperparameter sensitivity study characterizes only the Qwen3-4B / MMLU-Pro calibration cell and does not bound the potential benefit of per-cell tuning on other models or datasets. Prompt-level bootstrap intervals quantify prompt-sampling uncertainty for a fixed rollout and do not measure variability across independently generated rollouts. In addition, some numeric values in the loaded text (such as benchmark sizes, hyperparameter values, and some confidence intervals and tail probabilities) appear as placeholders, so exact experimental configurations should be checked against the original tables.
