Agensh: Scaling Organizational Intelligence to 1,024 Agents
Synopsis
Agensh introduces a self-organized multi-agent harness without a central orchestrator, in which concurrent workers run an asynchronous cooperation loop supported by a shared workspace, a message interface, and shared context; on the five hardest ProgramBench tasks it raises the mean final test-pass rate from 19.31% with 1 agent to 28.78% with 128 agents, and on pandoc from 33.89% with 1 agent to 55.06% with 1,024 agents.
Interpretation
Agensh removes the central orchestrator from the multi-agent harness and lets concurrent workers discover, claim, and complete sub-tasks themselves, using three infrastructure components — a shared workspace, a message interface, and shared context — to turn individual progress into organizational output. Compared with orchestrator-worker frameworks such as Codex sub-agent, Claude Code sub-agent and agent teams, Copilot fleet, and Kimi Agent Swarm, responsibility for task discovery, dependency coordination, and result integration is pushed down to the workers themselves. The paper specifies a five-step cooperation loop (gather context, claim sub-task, take action, verify results, merge progress) and the three infrastructure components, noting that the shared workspace is implemented with Gitea, the message interface with Mattermost, and shared context follows the core idea of DeLM with adapted tool formats and worker instructions.
On the five hardest ProgramBench tasks, agent count itself acts as a new scaling dimension: with the model GPT-5.6-sol (high), the underlying Copilot harness, and a 6-hour budget held fixed, mean final test-pass rates are 19.31%, 20.68%, 26.52%, and 28.78% for 1, 8, 32, and 128 agents. Earlier scaling discussions for multi-agent systems centered on model capability or orchestration strategy; here agent count is treated as an independently adjustable scaling axis, with a monotonic trend across five tasks. The five tasks are FFmpeg, gromacs, pandoc, PHP-src, and ctags, whose reference repositories contain thousands of files and hundreds of thousands to millions of lines of code; going from 1 to 128 agents yields 9.47 percentage points, an approximately 49% relative improvement.
More concurrent agents improve not only the final score but also the latency to reach a given score: on pandoc, 128 agents exceed a 30% test-pass rate at the 30-minute checkpoint, while 32 and 8 agents first exceed that threshold at the 60- and 90-minute checkpoints, and the single-agent run stays below it throughout the first two hours. This places latency and quality within the same scaling frame, showing that agent count can buy earlier convergence within a time budget. The evidence comes from checkpoint comparisons under the same 6-hour budget, with model, harness, and evaluation configuration held constant.
Scaling the organization to 1,024 agents on pandoc raises the final test-pass rate from 33.89% with 1 agent to 50.94% with 128 agents and 55.06% with 1,024 agents; worker trajectories show forms of self-organized cooperation emerging and standardizing as the organization grows. At 8 agents, workers agree on interfaces and resolve overlapping claims themselves; at 32 agents, multiple workers jointly contribute and a broader set of peers joins integration; at 128 agents, workers select reviewers based on prior experience, transfer integration responsibility, and reuse standardized integration protocols; at 1,024 agents, role specialization appears at organization scale and workers in the same technical area take over after a peer's attempt fails. All workers follow the same loop and receive the same prompt except for worker IDs, so the cooperation forms can be attributed to scale rather than individual instructions; the 1,024-worker experiment is distributed across 16 nodes with 64 agents per node.
Perspective
The results apply to the ProgramBench setting of rebuilding reference software behavior from scratch without Internet access under a fixed 6-hour budget, evaluated with GPT-5.6-sol (high) and Copilot as the underlying single-agent harness; the paper describes Agensh as an organization layer above a single-agent harness, so its gains apply to long-horizon engineering tasks that can be split into many parallel sub-tasks and coordinated through a shared workspace and message interface. For teams wanting to raise completion on complex tasks or shorten the time to a target score without swapping models, the trio of shared workspace, message interface, and shared context offers a directly reusable organizational structure.
The paper reports test-pass rates and trajectory observations on ProgramBench; whether the benefit curve of agent count remains similarly monotonic on other task families, other underlying models, or other underlying harnesses still needs more evidence. Shared context entries are capped in length (100 characters, 300 for PATCH_SUMMARY), so the effect of that compression over longer task horizons is worth watching. In addition, the 1,024-agent experiment is reported only on pandoc, leaving cross-task behavior at very large scale an open question.
