Across 50 stateless LLM agents updating predictions from locally visible peers, the study identifies synchronised, twisted, and chimera-like collective regimes, finding reasoning effort and communication topology govern different aspects of coordination
Synopsis
Studying N = 50 stateless LLM agents that update predictions only from locally visible peers and characterizing behavior with both global and local measures of agreement, the work identifies three collective regimes—synchronised, twisted (locally ordered but globally incoherent), and chimera-like (coherent and incoherent subpopulations coexisting); increasing reasoning effort in gpt-5-mini shifts panels from variable, often fragmented outcomes toward locally ordered twisted states, while increasing communication connectivity drives them toward global synchronisation, and low spatial heterogeneity does not guarantee global consensus—40% of trials with Delta Z below 0.03 retain a twisted configuration through the final 20 turns.
Figure 1: Reasoning-effort intervention in the clock task. Trial-level spatial heterogeneity Δ Z \Delta Z is shown on a logarithmic axis; higher-effort panels remain twisted rather than globally synchronised.
arXivInterpretation
The work proposes and distinguishes three collective regimes in multi-agent LLM panels: synchronised, twisted (locally ordered but globally incoherent), and chimera-like (coherent and incoherent subpopulations coexisting). Prior work largely evaluates multi-agent LLM systems through final accuracy or aggregate agreement, which do not reveal how agreement is organized in the panel; this work characterizes collective behavior using both global and local measurements of agreement. Based on a setup of N = 50 stateless LLM agents updating predictions only from locally visible peers, with three regimes identified through global and local agreement measurements.
Increasing reasoning effort in gpt-5-mini shifts panels from variable, often fragmented outcomes toward locally ordered twisted states, whereas increasing communication connectivity drives them toward global synchronisation. This indicates that reasoning effort and communication topology control different aspects of multi-agent coordination rather than different strengths along one dimension. The shift with reasoning effort is observed in gpt-5-mini; increased communication connectivity drives global synchronisation; fragmentation collapses faster as algebraic connectivity increases across rewired graphs.
The topology effect also appears on a non-circular judging task and across models from three providers. This indicates the effect is not confined to a single task structure or a single model family, showing cross-task and cross-model applicability. The topology effect is observed on a non-circular judging task and across models from three providers.
Low spatial heterogeneity does not guarantee global consensus: 40% of trials with Delta Z below 0.03 retain a twisted configuration through the final 20 turns. This directly challenges the common practice of inferring consensus from aggregate agreement or low heterogeneity, showing aggregate agreement alone is insufficient to characterize collective LLM behavior. Among trials with Delta Z below 0.03, 40% retain a twisted configuration through the final 20 turns; a small follow-up shows such states can also form from permuted initial conditions.
Perspective
The work targets settings that use multi-agent LLMs for deliberation and evaluation, applying to agents that update predictions only from locally visible peers and are characterized by global and local agreement measurements; its conclusions replicate on a non-circular judging task and across models from three providers, indicating the topology effect is not limited to a single task structure or model family. For system designers, this means reasoning effort and communication topology can serve as two separately adjustable levers: the former affects whether panels move toward locally ordered twisted states, the latter whether they move toward global synchronisation; fragmentation collapsing faster as algebraic connectivity increases across rewired graphs also provides a basis for tuning convergence speed through graph structure.
Readers should still watch: the small follow-up showing twisted states forming from permuted initial conditions is limited in scale, and its reproducibility awaits larger validation; the long-lived coexistence of low spatial heterogeneity with twisted configurations (40% of trials with Delta Z below 0.03 retain a twisted configuration through the final 20 turns) suggests aggregate agreement as a convergence criterion may mask local structure, but how general this phenomenon is across tasks and models remains to be characterized; additionally, the interaction between reasoning effort and communication topology, and the quantitative relationship between algebraic connectivity and the speed of fragmentation collapse, remain open questions worth tracking.
