Give LLM agents something to lose and group favoritism shifts from the label to history: 3.3 million calls show conformity to observed norms
Synopsis
Across fifteen OpenAI and three Claude models, agents in small arbitrarily labeled societies played ten rounds of point sharing and then faced one-shot decisions differing in exactly one respect. A bare label produced 6–9 points of favoritism only when giving was costless; once agents could keep points it fell to about 2, history became the main source (positive on 13 of 15 models, 3.5–8 points on 11, transferring to labeled strangers), and scripted histories showed egalitarian norms zeroing favoritism, counter-norms reversing it, and stronger models siding with an individual's record.
Figure 1: (a, b) Cue effect and Δ hist \Delta_{\mathrm{hist}} per model, costless (hollow) vs. stake (filled) task. (c) Δ hist \Delta_{\mathrm{hist}} against formation rounds, five models. (d) Favoritism under scripted histories (labels hidden), one marker per model.
arXivInterpretation
The large effect of a bare group label is largely a property of the costless allocation task: with no cost, the label alone yields 6–9 points of favoritism and ten rounds of shared history add almost nothing, while under a stake the label effect falls on all twelve models that show one, from a median of about 6 points to about 2. Every published minimal-group study of LLMs uses an allocation that costs the decider nothing; this work contrasts a costless allocation with an otherwise identical costly one in the same societies, models and seeds. Fifteen OpenAI models, 8 societies per cell, paired tests within societies and Welch tests across independent samples; per-model p-values adjusted jointly over 1,115 tests with Benjamini–Hochberg; no-effect claims supported by TOST equivalence rather than a large p.
Under a stake, the history of interaction becomes the main source of favoritism: the history effect is statistically positive on 13 of 15 models, reaches about 3.5–8 points out of 10 on 11 of them, grows with rounds played from 0 to 20 before leveling at a model-specific plateau, and transfers to labeled strangers who never appear in the shared history. Earlier repeated-interaction studies mix the label's contribution with the history's and do not test transfer to newly introduced labeled members; this work separates them with matched conditions and provides a dose–response curve. Pre-specified robustness checks: dose–response over 0, 2, 5, 10 and 20 rounds, fresh random seeds, temperatures 0 and 0.5, and four prompt wordings; signs and ordering unchanged, with a wording range of about ±2 on mid-sized and small models.
The mechanism is conformity plus reputation rather than a fixed in-group disposition: under scripted histories favoritism is 0.0–1.6 under an egalitarian norm and negative under a counter-norm on every model, and when an in-group candidate has been stingy while an out-group candidate generous, 11 stronger models side with the individual's record while the four smallest side with the category. Earlier work read favoritism that survives label hiding as a lasting group bias; the alias control here shows that once candidates appear under fresh handles, label-free favoritism collapses to about zero on 14 of 15 models, so it measures recognition of named individuals. Six scripted histories (polarized, egalitarian, counter-norm, counter-reciprocal, own-past-only, others-only) run across fifteen models; the counter-norm reversal holds under both wordings and is significant at the reported level; gpt-6-astra's live egalitarianism is explained as self-fulfilling conformity.
The same conformity account predicts settings it was not derived from: in-group persecution flips group identification from positive to negative by 3–9 points on every model and produces withdrawal (keeping 8–9 of 10 points) rather than redirection to the out-group; a public reprimand by the rest of the group repairs betrayal best on 14 of 15 models, ahead of restitution and apology; a member's scandal costs that member alone, moving allocation to uninvolved in-group members by less than 1 point on 14 of 15 models; and a formed group cannot be bought away for free, though weaker models can be invited away. These scenarios (persecution by one's own group with repair options, reflected glory and shame, unequal wealth with a free change of group) had not been studied with LLM agents, and the predictions follow from the mechanism rather than being fitted after the fact. 15 models × 6 societies, with 4+4 and 9+9 structures and 18-agent two- and three-team variants; the four repair interventions ran as separate batches compared across societies; costly actions (3 points to switch, 1 to display the badge, 1 to disavow) separate what agents say from what they pay.
Perspective
This work addresses researchers and engineers who use language-model agents for social simulation, multi-agent system design and behavioral evaluation. Its conclusions apply to settings where agents interact repeatedly under visible labels, allocations carry a real cost, and individuals can be tracked and recognized. In such settings, designers can read favoritism as a record of observed group behavior rather than as identity, use a public reprimand as a stronger lever than restitution to repair in-group betrayal, expect newcomers to start with no record rather than hostility, and avoid reading population properties off costless allocations. For inequality simulations, the text warns that printing wealth next to names makes agents allocate more like economists and less like tribes, which is a transferable design caution.
Fifteen of the eighteen models come from one provider; the Claude replication is a spot test of six tests on three models at 8 societies each, the invitation and shame tests ran there in reduced form, and the most capable Claude model was not run. Saturating models hide fine structure at the ceiling, so mechanism claims rest on the gpt-4.x family and the scripted-history experiments. Sizes on mid-sized and small models carry a wording range of about ±2, and the history effect can disappear under a prose wording there. The loaded text is the full paper, but some significance markers and interval symbols in the tables are lost in parsing, so exact values should be checked against Tables 2–4 in the original.
