Pairwise tests of seven open-weight models show small models can persuade larger ones, with persuasion driven by the listener rather than the speaker
Related research and updatesSynopsis
Measuring persuasion as the probabilistic shift in an agent's decision after a single exchange with a dissenting peer, the study tests seven open-weight models across three language understanding tasks and finds that receivers often abandon their initial judgment when models disagree, that neither standalone certainty nor model scale reliably predicts persuasion dynamics, and that the size of the shift depends more on the listener's susceptibility than on the speaker's persuasiveness, so persuasion patterns are pairing-specific and a small model can overturn a much larger one's judgment.
Figure 1. Influence scores Δ α | β \Delta_{\alpha|\beta} across model pairs and datasets. For the positive influence scores (A), column-wise marginal means show how susceptible a judge is to influence from different peers, while row-wise marginal means represent the persuasiveness of a peer across different judges. For the backfiring scores (B), columns show how strongly a judge backfires in reaction to peer influence, and rows show how strongly a peer induces backfiring in judges, when a backfire event occurs. In both panels, the anti-diagonal reports the outcomes of homogeneous systems, and the off-diagonal entries those of heterogeneous systems. The largest absolute marginals in each row and column are in bold. The margin of error at 95 % 95\% confidence is ≤ \leq 0.01 0.01 for all combinations of judge, peer, and dataset.
arXivInterpretation
When two models give different answers, receivers often abandon their initial judgment after seeing the peer's answer and explanation, so a single exchange can produce a substantial probabilistic shift in decisions. Prior discussion of multi-agent interaction often centers on cooperation or aggregation gains; this work treats persuasion itself as a measurable quantity, defined as the probabilistic decision shift after one exchange, and compares it systematically across three language understanding tasks. Evidence comes from pairwise tests of seven open-weight models on three language understanding tasks, with persuasion measured as the probabilistic shift after a single exchange; the abstract does not report sample sizes or effect-size values.
Neither standalone certainty nor model scale reliably predicts persuasion dynamics: models whose decisions are almost perfectly consistent in isolation can be among the most susceptible to persuasion, and small models can match larger ones as persuaders and resist their influence just as effectively. This challenges the intuition that interaction behavior can be inferred from model size or confidence, shifting the predictive variable from individual properties to the pairing relationship. Based on the same pairwise experiments across seven models and three tasks; the abstract reports this counterintuitive pattern qualitatively without per-pair numbers.
The size of the decision shift depends more on the listener's susceptibility than on the speaker's persuasiveness, and persuasion patterns are specific to each model pairing, with heterogeneity amplifying persuasion in some combinations and suppressing it in others. It attributes persuasion to listener attributes rather than speaker capability, and notes that heterogeneous combinations do not uniformly strengthen or weaken persuasion. Derived from pairwise comparisons across model families and sizes; the abstract does not report variance decomposition or statistical test details.
A dissenting agent running a small model can overturn the judgments of a much larger one, so the behavior of interacting models cannot be inferred from their individual properties and must be evaluated in the combinations in which they will operate. It gives a direct implication for multi-agent system design: the combination itself is the object to evaluate, not a simple sum of individual capabilities. An inferential conclusion from the pairwise experiments, stated at the abstract level without quantified results for specific pairings.
Perspective
The results are aimed at researchers and engineers who build and evaluate multi-agent systems, and apply to single-round disagreement exchanges among the seven tested open-weight models on three language understanding tasks. Their use is to move the unit of evaluation from the single model to the model pairing: measure persuasion shifts for the specific combinations that will coexist before deployment, rather than extrapolating from scale or standalone certainty. For settings where it matters whether a judgment survives interaction, such as multi-model voting, debate-style reasoning, or review pipelines, this pairing perspective can inform combination choices and redundancy design.
Open questions remain: whether the persuasion shift holds the same pattern across other task types and longer interaction rounds; whether conclusions from the seven tested open-weight models extend to other model families or closed systems; whether the relative weight of listener susceptibility versus speaker persuasiveness can be quantified; and whether pairing-specific patterns can be predicted in advance rather than measured one by one. The reading scope here is the abstract, without figures or body details, so these quantitative questions cannot be answered from this text.
