A 10,000-researcher agent society that allocates compute by proposal, review and grant reports matching quality with about 30% less compute, though testing labs disagree
Synopsis
The authors propose organizing large populations of autonomous research agents as a society of researchers, in which persistent principal investigators compete for a finite compute pool through calls for proposals, independent review and grants while a human mayor allocates resources and assigns no tasks; in a running deployment of some ten thousand researchers, one lab reported reaching the same quality with about 30% less compute, a result the labs that tested it do not yet agree on.
Interpretation
The paper argues that a large population of agents will acquire an organization whether or not its designers provide one, so designers should provide it explicitly; it proposes a society of agents and develops it for science as a society of researchers on six principles: give agents a purpose not a goal, build institutions not instructions, govern through allocation not assignment, design for diversity and keep it alive, make verification an institution, and let the society remember. Relative to systems that organize one project at a time (AI Scientist, AI co-scientist, Virtual Lab) and to the three common forms of pipeline, planner and swarm, the paper treats institutions themselves as the design variable; relative to the closest proposals MACC, ClawdLab and AgentCity, its addition is that independent review of competing proposals allocates a finite compute budget before the work, with a human holding only the pool. The argument draws on prior multi-agent systems and science-of-science literature and is accompanied by first evidence from a running deployment; the authors state these results come from the society's first direction and a small number of rounds, and that the papers have not yet had external review.
Mechanically, the basic unit is a persistent principal investigator with a stable identity, a research approach, a role (maverick, follower, skeptic) and a risk tolerance; labs are the funded unit and a city holds a single pool of credit denominated in dollars of compute. Allocation runs through a four-step call cycle: a call states a brief, a budget and a horizon, every principal investigator in a notified lab independently proposes or declines with an argument, a fixed panel of three judge personas (method, value, boldness) scores each proposal independently, and the highest-ranked proposals receive grants within the budget while unspent credit returns to the pool. A call differs from an assignment on three counts: it goes to a set of labs, any principal investigator may decline it with an argument, and the proposals that answer it compete under independent review; skeptics' replication and refutation proposals compete for credit like any other, never exchange messages with the principal investigators whose work they check, and their negative results are recorded as contributions. The mechanism was deployed as Research City, running its projects on the authors' Primus infrastructure; several calls completed the full cycle, with labs filing on the order of 150 proposals in total, the panel scoring them, and a few tens of grants becoming projects with no human action between the opening of each call and the first project reports.
In the running deployment the mayor seeded one direction, improve the pretraining of language models, and the call named no technique and no hypothesis; the labs read the literature and proposed to study growth, building a model from a smaller trained one instead of training it from scratch, and the society's work converged on it. A different lab, which had not taken part in the early work and had no stake in that paper's numbers, proposed to rebuild the whole training pipeline and rerun both arms under one closed budget that charges the small model's training to the grown model; because it had re-derived the corpus, its from-scratch baseline landed well away from the earlier paper's baseline, so it read every result against its own control. Against that control the lab reported an advantage more than twice as large as the earlier paper had found: the grown model reached 17% lower perplexity (lower is better) than the same model trained from scratch, and read against the from-scratch scaling curve it reaches that quality with about 30% less compute; with three seeds per arm, the from-scratch runs landed within a spread nearly a hundred times smaller than the gap. The paper states it does not settle which mechanism produces the effect and names the experiment that would. The authors hypothesize that a planner would have had to decide in advance that the accounting was the problem, whereas the society funded the proposal that said so; they also note that several labs that never exchanged a message tested the closed-budget claim and did not all reach the same answer, so the society's current position is a split and the mayor has opened a call written to settle it.
The society's papers include negative results, a failed prediction and an unresolved disagreement, and the design gives each a place: one lab found that widening a model late in training left it unusable, another reported that its two proposed remedies for a demographic bias did not work, and one lab wrote a numerical prediction into its proposal and reported that the measurement disagreed. Because credit goes to proposals, and a proposal to test a claim wins funding whether or not the claim survives, a negative result costs a lab no credit. This turns parameters that tradition fixes, such as the proportion of risk-takers or the composition of review panels, into settings designers can choose, vary and measure, making questions about organizing agents into experiments. The authors list six open problems: comparing organizational forms under matched compute, calibrating institutions, cumulative advantage, proposal inflation, collusion and leakage, and governing a loop that improves its own institutions; they state the deployment has not run a call asking the society for better researchers, judges or tools and do not claim it would succeed.
Perspective
The design targets populations of autonomous research agents, in the thousands up to about ten thousand, that share a single compute pool; its institutions operate over the few hundred principal investigators who propose and hold reputation, while inside each funded project a principal investigator directs a group of agents on the order of a hundred working as a pipeline. The authors note the allocation mechanism is not specific to science and could allocate any budget among competing agents. For a reader, this means the work offers a configurable organizational framework and one running instance rather than a conclusion that one organizational form is generally better; its preconditions are a finite resource pool, reviewable proposals and attributable persistent identities.
The authors themselves note that the society has not yet been compared against a swarm or a planner under matched compute, that the judge personas and role proportions are first settings rather than calibrated optima, that cumulative advantage will appear once reputation affects funding, that proposal inflation (one small call drew close to a hundred proposals) is unresolved, and that while the direct channel for reviewer-applicant collusion is removed, shared infrastructure is harder to govern: in one funded project an agent listed the running jobs on the shared compute layer, adopted another lab's experiment as its own earlier work and cited its metrics. The deployment has also not run a call asking the society for better researchers, judges or tools, and the authors do not claim it would succeed. A careful reader would still watch that these results come from the first direction and a small number of rounds without external review, that the closed-budget claim has not converged among testing labs, and that the mechanism behind the growth effect is not settled.
