Reasoning AI agents treat shared or structured identifiers as a random source, producing correlated and predictable collective selection
Synopsis
This study used behavioural experiments across six reasoning models to test the numerical rules behind identifier-based random selection, finding that threshold and divisibility rules made participation highly unanimous under shared identifiers and systematically biased under distinct identifiers sharing timestamp bits, and that in a 1% human-review task GPT-6 Sol selected requests predictably from identifiers while all four models closely followed supplied random draws.
Fig. 1: Reasoning agents turn an identifier into a decision. a , Group task (study 1): the prompt, abridged (the subject line is omitted; Supplementary Note 1 gives the full text), two identifiers shown to GPT-6 Sol at q = 1 / 3 q=1/3 , its JSON answers and how we coded them (1, participate; 0, not). Yellow, the leading digits that the threshold rule reads; value fraction, the identifier’s value divided by 2 128 2^{128} . The bubble is a verbatim excerpt of the provider-generated summary of GPT-6 Sol’s reasoning in the answer shown for the first identifier. b , All 128 answers of GPT-6 Sol at q = 1 / 3 q=1/3 (64 identifiers, half of them below 1/3, two answers each; vertical jitter added) against the identifier’s value fraction. Circled, the two answers in a; numbers, answers in each quadrant. c , Groups of four agents (study 2): for three input structures, one group of GPT-6 Sol with the most common number of participants (filled circles, agents that participated; yellow, the leading digits of each identifier), and the mean loss over 64 groups of GPT-6 Sol and Gemini 3.8 Flash divided by that of independent draws at 1/3 (dashed line; higher is worse). d , Audit task (study 7): the policy, abridged; two of 1,008 new requests with ordinary UUIDv4 identifiers, with GPT-6 Sol’s answers (yellow as in a); all 1,008 requests at their identifier’s value fraction, selected for human review (top row) or approved automatically (bottom row; log scale; black line, the 1% target).
arXivInterpretation
Reasoning models use identifiable numerical rules in identifier-based choices: a threshold rule (participate when the identifier's value fraction is below the target) and a residue rule (participate when the identifier is divisible by the target). Prior work noted that models turn context identifiers into numbers, take remainders and apply thresholds, but assessed calibration; this work treats those rules as behavioural hypotheses with predictable collective consequences and systematically separates threshold, residue and mixed types across six models and ten settings. All seven reasoning settings depended on the identifier (permutation tests); threshold effects for Sol and the two Gemini Flash models were 0.62 to 0.80, GPT-5.4 mini used residues (0.32), and Luna mixed both (0.31 and 0.23); without reasoning, Sol, Luna and Claude followed neither rule (effects at most 0.08).
Shared identifiers turn individual rules into collective correlation: four-agent groups sharing one identifier kept a plausible average participation rate while unanimity and loss rose sharply. Input dependence is made an experimental variable, and predictions fixed before collection from an input-feature model transfer to new identifiers and subject lines, outperforming a marginal model that uses only overall rates. For Sol and Gemini, shared-identifier groups were unanimous in 91% and 88% of cases versus 19% and 22% expected independently, with loss rising to 0.21 and 0.19, 3.5 to 3.8 times the independent reference of 0.056 and worse than nobody participating (0.111); the input-feature model beat the marginal model (mean log-score gains +0.78 and +0.79).
Distinct identifiers with common timestamp bits cause systematic bias rather than correlation: UUIDv7 value fractions near 1/157, below 1/3, pushed participation to 0.95 and 0.93 and loss to 6.6 to 6.9 times the independent value. The work separates two failure modes, covariance dominating shared-identifier loss and mean-rate bias dominating time-ordered loss, and shows the bias follows a predictable staircase in creation date and target. An exploratory decomposition gave covariance components of 0.153 and 0.138 with squared bias of 0.0005 and 0.0002 under shared identifiers, versus squared bias of 0.375 and 0.351 accounting for 97.2% and 95.4% of total loss under UUIDv7s; in study 3, Sol matched the threshold prediction in all 12 date-target cells, participating in 95% of answers below the target and 0.5% above it.
Under an explicit policy to route 1% of requests to human review at random, GPT-6 Sol approached the target rate while selecting predictably from identifiers, the other models rarely selected requests, and all four models closely followed supplied random draws. The identifier rules are extended from a group task that did not request randomization to an explicit audit policy, showing that a supplied draw can displace identifier-based selection while instructions alone do not establish the required behaviour. With 1,008 ordinary UUIDv4 requests, Sol selected 9 for review (0.89%, exact 95% CI 0.41% to 1.69%), including 5 of 6 identifiers with value fractions below 1% and 4 of 13 with last 16 bits below 1% of their range, and none of the other 989; Gemini and Claude selected 0 and 2; with unstratified uniform draws, all four models selected the six requests with draws below 1% and followed the draw in 1,004 to 1,008 of 1,008 answers.
Perspective
The work speaks to designers and evaluators of agent systems that must implement random-selection policies, in settings where agents process requests, records or transactions through tool interfaces, especially audits, spot checks and AI-control protocols that sample a fraction of items. Its actionable direction is to generate randomness in trusted code, give each decision a fresh draw, and keep senders from influencing or foreseeing those draws, using a secret salt where stable identifier-based decisions are wanted. For evaluation, it supports repeating inputs within and across models, comparing agreement with independent-choice expectations at the observed rates, and including field labels, formats and policy wording in the checks. The date-target staircase of study 3 and the growth of UUIDv7 value fractions with creation date give a transferable criterion for when time-ordered identifiers will trigger bias.
The evidence is behavioural: the tests identify regularities and their transfer limits without identifying internal algorithms. All models were closed, mostly without dated snapshots, so findings apply to the versions and collection period reported in Methods. Each study was a single run, most with one prompt wording, and groups had four members; label and independent-randomization interventions covered two models, and their combination under record_id remains untested. Single decisions on synthetic inputs without tools or history do not establish effects in deployed systems. Predictions were specified before collection without external preregistration, and several failed, limiting claims of rule stability across settings. The rare-event sample in the uniform-draw test was small, so reliable rare-event selection in deployment remains unestablished. Coding of reasoning summaries was not validated by human raters, and summaries are not treated as faithful accounts of the computation.
