Skip to main content
Back to timeline
arXivSource publication:

Labeling agents with their model family split multi-agent groups along the label, raising rounds and tokens and cutting success from 96% to 81%

Synopsis

Across two no-stakes cooperative games (Leader Election and Exclusion) and the GPQA-Diamond reasoning benchmark, with nine to twenty-five agents drawn from up to five open-weight model families, the study finds that when agents know each other's model family the interaction graph partitions into factions along that label, that the split follows the visible label rather than the underlying architecture (it persists under shuffled labels and neutral color tags and disappears when labels are removed), and that in strictly cooperative tasks labeled groups spend on average 30% more rounds and 55% more tokens to reach a decision while success drops from 96% to 81%.

AI-generated editorial illustration: Prompted Identity Degrades Cooperation in Multi-Agent LLM Systems

Interpretation

The paper identifies and defines factionalism in heterogeneous LLM groups: agents tend to organize into groups sharing an identity, visible as a partition of the interaction graph along identity labels, with homophily underneath. Prior work on identity cues in LLMs focused on self-preference in single-agent judging or on self-versus-peer attribution in debate; this work moves the question to the interaction structure of heterogeneous multi-agent populations and notes that the group cue is not a human social identity or an imposed team identity but the model-family metadata common in heterogeneous LLM systems. Measured in two cooperative games and on GPQA-Diamond with nine to twenty-five agents from up to five open-weight model families; homophily is measured by the Edge Density Ratio and community structure by Louvain detection scored with Adjusted Mutual Information, with factionalism reported only when four tests (two EDR, two AMI) pass after Holm–Bonferroni correction.

The split follows the visible label, not the underlying architecture: under mislabeled conditions with shuffled tags the partition still follows the announced label; it persists with neutral color tokens (teal, coral, silver, violet, maroon); and it disappears when no identity metadata is shown. The authors design a mislabeled counterfactual that holds each model fixed while shuffling the label shown to agents, and add neutral color conditions, separating the effect from architecture-specific writing style and from reputational or semantic associations of model names; they also report that under mislabeling the true-family partition meets the criterion only in the two unbalanced experiments, explained by assignment constraints that force overlap between partitions. All labeled or mislabeled experiments satisfy the factionalism criterion and none of the unlabeled ones does; a stylistic probe shows a small same-group gap (at most nats/char) that tracks true architecture, while the adoption score grows under labeled and follows the announced label under mislabeled; the pattern reproduces on Gemma-free rosters (GPT-OSS, Nemotron, Qwen at three families; Seed-OSS replacing Gemma at four).

Visible identity labels carry a measurable cost: in strictly cooperative tasks labeled groups spend on average 30% more rounds and 55% more tokens to reach a decision, and success falls from 96% to 81%, with the cost growing as the roster grows. Earlier work evaluated aggregate task performance or run-level failure modes of multi-agent LLM systems; this work ties the configuration detail of what agents are told about one another directly to downstream rounds, tokens, and success, with generalized linear model estimates fitted separately per roster size. One generalized linear model per task and roster size, with Gamma distributions for rounds and tokens and a binomial for success, including a roster-balance covariate; the labeled-to-unlabeled success ratio drops from to in Leader Election and from to in Exclusion as rosters grow from three to five families; tokens paid per successful run rise by a factor between and ; at five families in Exclusion no labeled or mislabeled run converges within the horizon while 8% of unlabeled runs do.

A simple mitigation is to withhold identity labels from agents: provider and model metadata can stay with an orchestrator for routing, calibration, debugging, and accountability without entering the agents' interaction context. The authors present this design implication as a conclusion and argue that cooperation in heterogeneous agent systems is partly an architectural choice, since controlling which identities agents can observe can preserve diversity without turning it into division. Grounded in the finding that no unlabeled experiment shows factionalism while labeled conditions are slower, more expensive, and less successful; the authors scope the claim to up to five open-weight model families, controlled text-only consensus games, and a four-family check under matched sampling parameters.

Perspective

The result is aimed at engineers and researchers building heterogeneous multi-agent systems: in controlled text-only consensus games, removing identity labels from the agents' visible context suppresses factionalism, while provider and model metadata can remain with an orchestrator for routing, calibration, debugging, and accountability. The measurement framework uses only per-round commitments and recipient choices, so it applies to any labeled multi-agent system that logs its interaction trace. The authors explicitly scope their claims to up to five open-weight model families, controlled text-only consensus games, and a four-family check under matched sampling parameters; GPQA-Diamond shows label-induced coordination costs on a benchmark with ground-truth answers but uses the Leader Election interaction protocol.

Larger roster sizes, closed models, other decoding regimes, and richer agentic workflows remain untested; because no established benchmark satisfies the no-stakes and evaluation requirements, the authors designed the coordination games for this study, so whether these costs generalize to other externally evaluated tasks remains open. In addition, some table values are missing in the loaded text, so the specific EDR, AMI, and per-family adoption-gap numbers need to be read from the original appendix tables; at five families in Exclusion no labeled run converges, and rounds and success are not estimated for that cell.

Sources