Skip to main content
Back to timeline
arXivSource publication:

Multi-agent systems homogenize in code generation, hiring, and peer review, and neither sampling stochasticity nor mixed-model configurations mitigates the risk

Synopsis

The study operationalizes homogenization in multi-agent systems through conformity, polarization, and inertia, evaluates it on code generation, hiring, and peer review, and finds that agent interactions reduce diversity and amplify shared failures, while sampling stochasticity and mixed-model configurations fail to mitigate these risks.

Source-provided article image: Homogenization in Multi-Agent Systems
Figure 1 ·

Figure 1 : Homogenization in MAS and the resulting failure modes.

arXiv

Interpretation

The paper introduces and demonstrates homogenization as a failure mode of multi-agent systems, characterized by three metrics: conformity (declining variance of agent scores across rounds), polarization (drift of the system mean toward scale endpoints), and inertia (declining rate of change of the mean across rounds). Prior work on homogenization largely treated models as isolated decision makers or studied subjective debate and games; this work operationalizes homogenization and extends it to real-world multi-agent applications, with inertia newly introduced by the authors. The three metrics are defined for orchestrator-driven multi-agent systems and measured across three tasks, multiple models, and mixed configurations, reported as means with confidence intervals over independent random-seed trials.

In code generation, multi-agent systems reduce overall security vulnerability rates but homogenization amplifies correlated errors, creating systemic security blind spots that are shared across model families. Prior evaluations emphasized aggregate performance gains; this work disaggregates by vulnerability type and conditions on deployability, showing that specific vulnerabilities (e.g., py/url-redirection, py/log-injection) persist or amplify across rounds. Evaluated on the CodeLMSec dataset of 280 Python and C++ prompts, with vulnerabilities detected using CodeQL security query suites and filtered by the reviewer panel's average deployability score to isolate flaws that evade the entire system.

In hiring, stereotypical bias injected into a single agent spreads rapidly to the whole evaluator panel and persists even after the biased agent is removed, indicating heightened susceptibility to manipulation. Prior work largely measured static bias; this work injects adversarial instructions for exactly two rounds to show cross-agent transmission and residual bias after removal. Uses a synthetic counterfactual résumé dataset constructed following Tan et al. (2026), covering gender and ethnicity variants, with each counterfactual variant evaluated independently and demographic bias quantified as systematic score disparity.

In peer review, homogenization drives evaluations toward convergent preferences, widening the score gap between research areas across interaction rounds. Prior evaluations of automated review focused on agreement with human reviewers; this work shows that multi-agent interaction hardens latent model preferences into systematic cross-discipline disparities. Adopts the multi-agent structure of Sahu et al. (2025) with three reviewers, a meta-reviewer, and an author, sampling ICLR 2025 submissions stratified by author-selected primary area and aggregating mean reviewer scores and distribution spread by area.

Perspective

The work targets researchers and deployers of orchestrator-driven multi-agent systems, applied to code generation, hiring, and peer review, three tasks that naturally involve multi-role collaboration. Its value lies in a reusable three-metric basis for auditing interaction dynamics before deployment rather than relying on aggregate performance alone; for teams planning to put multi-agent review or screening pipelines into practice, the results suggest monitoring cross-round score variance, mean drift, and rate of change.

Readers should note that homogenization is measured on agents' numerical output scores, while natural-language homogenization is not quantified; experiments are limited to orchestrator-driven multi-agent architectures, leaving other architectures open; hiring evaluation uses synthetic counterfactual résumés whose demographic signals were validated as recognizable by Gemini 3.1 Pro, so the gap to real résumés warrants further observation; peer review samples come from a specific stratified sample of ICLR 2025. In addition, several specific numbers in the main text (such as the magnitude of vulnerability reduction, polarization points, and round-to-round change) are omitted in the provided text, so exact effect sizes should be checked in the original figures.

Sources