BaRe-Mem uses Bayesian reliability memory to shift a central model toward autonomous reasoning as advice turns misleading, and finds capable workers earlier on MuSiQue
Synopsis
The work introduces BaRe-Mem, an online Bayesian reliability memory that estimates advisor reliability conditioned on the central model's internal belief representations, uses those estimates to modulate attention and to decide whether to consult; across nine benchmarks and six central models it is more robust to misleading advisor information than debate and majority voting, stays above autonomous reasoning on the more challenging tasks across all tested misleading levels, and in MuSiQue agent-team routing identifies capable workers earlier than routing by historical success counts.
Interpretation
BaRe-Mem models advisor reliability as Bayesian linear regression conditioned on the central model's internal belief representations, maintained by rank-one online updates with a Kalman gain, so usable reliability estimates emerge from sparse verified feedback. Existing memory approaches store textual trajectories and summaries or persistent latent states, but do not explicitly model advisor reliability conditioned on the central model's internal belief about the current question; BaRe-Mem lets these estimates scale with an expanding task stream and adapt to changes in advisor reliability. The paper provides posterior derivations, a minimum-mean-squared-error proof for the Kalman gain, and proofs of the rank-one update and order independence; experiments show clear gains over the no-feedback setting in the magnified region below roughly 170 samples, with performance then largely saturating.
Reliability estimates steer attention: for each advisor-response token, the unnormalized attention weight is rescaled before softmax by relative reliability, leaving the most reliable advisor unchanged while progressively downweighting less reliable ones. This mechanism introduces no trainable parameters and no additional training, and depends only on relative reliability among advisors, writing 'whom to trust' into attention rather than into prompts or parameter updates. The 'Advisors + memory' ablation, which keeps reliability-guided attention but removes the autonomous option, is often sufficient to maintain stable performance in the capability-supported regime, indicating that historical reliability estimates do reduce consultation's sensitivity to misleading advisor responses.
BaRe-Mem further compares estimated consultation ability with autonomous ability to decide per question whether to consult; in the capability-challenging regime the consultation ratio drops substantially as the misleading ratio rises while accuracy stays above the no-consultation baseline. Relative advisor weighting alone can determine whom to trust more but not whether the advisor pool as a whole is worth consulting; BaRe-Mem adds this missing gap, letting the central model rely on itself when external evidence is collectively unreliable. In Table 1, under the capability-challenging regime the consultation ratio falls from 90% to 21% for Qwen3-14B and from 85% to 18% for Phi-4, while BaRe-Mem remains above the 'No consultation' baseline across all tested misleading ratios; the predicted consultation gain and the real post-verification gain shift downward together as misleading information increases, with curves crossing zero close to a predicted gain of zero.
The same reliability memory extends to worker routing in agent teams: on MuSiQue's 2,417 tasks and 6,404 sub-tasks, BaRe-Mem achieves the highest task completion across both lead agents and all verification settings, and reaches higher completion with fewer worker calls. Compared with routing by historical success counts, BaRe-Mem conditions worker reliability on the current sub-task rather than using only an aggregate source-level success rate, which makes its advantage largest at small worker budgets. All routing strategies approach the same ceiling once every worker has been tried, so the difference lies in how many worker calls are needed to reach it; under dataset-verifier checking, task completion exceeds historical-success-count routing by several points, and the advantage persists across no check, lead-agent check, and dataset-verifier check.
Perspective
The results target deployment settings that need multi-agent consultation and team routing: a central model with access to internal belief representations, intermittent verified correctness feedback, and a long-horizon task stream. They apply where capabilities are heterogeneous and advisor reliability is task-dependent, and to agent teams that assign sub-tasks to workers; the paper uses MuSiQue-provided sub-tasks to isolate routing effects, so the routing benefit applies where task decomposition quality is held fixed. For practitioners seeking to reduce reliance on costly verification, the sparse-feedback experiments indicate that a small number of verified interactions already yields usable reliability estimates.
Reliability estimates remain above empirical accuracy at high misleading ratios: the paper notes that at 100% misleading the empirical accuracy is nearly 0 while estimates remain around 0.2, attributing this to residual unit observation noise in the predictive mapping and calling it 'an illustrative limit rather than a universal lower bound.' Readers should therefore watch how estimate calibration and the decision threshold behave in extreme misleading regimes. In addition, the team-routing experiments use dataset-provided sub-tasks to isolate routing effects, so the influence of task decomposition quality itself is not examined in that setting; the effect of report checking also varies by lead model, with the paper reporting that lead-agent checking improves over no check for five of six lead models while Ministral-8B behaves differently. These are areas for continued observation.
