Dynamic concurrency in three frontier coding agents: limited long-horizon gains, token use rising to 1.41–3.31 times sequential
Synopsis
This work presents the first systematic study of native dynamic concurrency in Codex, Claude Code, and Kimi Code, running controlled concurrent-versus-sequential comparisons on 354 tasks from four long-horizon benchmarks and one bounded benchmark for 2,124 executions; concurrency helps most on long-horizon tasks and little or negatively on bounded ones, raises token use to 1.41–3.31 times sequential, and yields a taxonomy of four main categories, 13 subcategories, and 28 concurrency failure patterns.
Figure 1. Overview of the dynamic concurrency mechanisms used by the three coding agents in our study. Separate dashed panels show main agents on the left and sub-agents on the right across three tool rows. Claude executes a scripted sequence of parallel stages and joins. Kimi launches a batch of independent sub-agents and waits for one ordered report. Codex coordinates direct and nested sub-agents while the main agent continues working and integrates returned results.
arXivInterpretation
Dynamic concurrency does not consistently improve task success: resolution effects range from an 11.9 percentage-point decrease to a 2.3-point increase, with gains concentrated on long-horizon benchmarks, especially the longest-horizon LoopsBench, where Codex and Claude Code improve task pass rates by up to 14.3 percentage points, while the bounded SWE-bench Verified shows no gain and substantial decreases for Claude Code and Kimi Code. Prior evaluations focused on single-agent systems or multi-agent systems with predefined collaboration protocols, and predominantly on bounded tasks such as SWE-bench with small patches; this work instead evaluates the native dynamic concurrency of frontier coding agents on tasks requiring tens of thousands of lines of code. 354 tasks, three agents, two settings, 2,124 executions, 11.2 billion recorded tokens corresponding to $20,780.02 at standard API rates; exact McNemar tests with Holm correction on pooled task groups show the overall declines for Claude Code and Kimi Code are statistically significant, while the improvement for Codex is not.
Concurrency produces unique successes: concurrent mode resolves tasks that sequential mode does not across almost all agents and benchmarks, for example Claude Code solves four LoopsBench tasks exclusively under concurrency and five SWE-bench Verified tasks exclusively under concurrency despite a lower overall pass rate, while Codex achieves identical aggregate performance but uniquely solves nine tasks in each mode. Aggregate pass rates alone hide this complementarity; the overlap analysis shows concurrency expands the capability set of coding agents rather than only shifting the average. Based on overlap of successful task sets between the two modes per agent; because no sampled ProgramBench task is fully resolved by any agent in either mode, a ProgramBench task counts as successful if its test pass rate is above zero.
Concurrency substantially raises resource consumption without reliably reducing time: mean shell calls increase for every agent and benchmark, token consumption reaches 1.41–3.31 times sequential, token use per solved task increases for all three agents, with Codex rising from 7.69M to 24.12M, and mean runtime increases in 14 of the 15 agent-benchmark combinations. This quantifies the benefit-cost trade-off, indicating that current concurrency mechanisms generally trade substantially greater resource consumption for limited aggregate effectiveness gains. Based on recorded provider usage; among the 366 cases solved in both modes, concurrent execution finishes faster in 102 cases (27.9%), and in 46 cases (12.6%) it is at least 20% faster.
Trajectory analysis yields a taxonomy of four main categories, 13 subcategories, and 28 concurrency failure patterns covering task orchestration, execution governance, global context management, and shared state and merge; these occur in 650 reviewed concurrent trajectories with 804 pattern instances, with shared state and merge the largest category (33.2%) and concurrent writes at 26.5%. Prior work analyzed single-agent failures or coordination failures under predefined protocols; this taxonomy specifically characterizes coordination failures of dynamic concurrency and distinguishes functional failures from performance failures under the available budget. All 1,062 concurrent trajectories were manually annotated; a codebook was built from a 50-trajectory open-coding pilot, each case was independently annotated by two of three annotators with recorded reasoning, disagreements were discussed and adjudicated by a senior author, and a single Cohen's kappa agreement value is reported.
Perspective
The study covers three frontier coding agents with their native scaffolds and highest reasoning effort (Codex with GPT-5.4, Claude Code with Claude Opus 5, Kimi Code with Kimi K3), so the findings apply within the tested agents and benchmarks. For practitioners, dynamic concurrency is more promising for difficult, long-horizon tasks that can be clearly decomposed into independent components, support multiple parallel attempts at the same problem, or allow independent sub-agents to validate the main implementation, and it should be enabled selectively rather than by default. For researchers, the taxonomy points to task decomposition and ownership assignment, runtime monitoring and intervention, communicating requirements and interface contracts across isolated contexts, and isolation, synchronization, and conflict detection for shared files and environments as directions to pursue. The authors release the 2,124-cell manifest, evaluator summaries, runner and parser versions, trajectory hashes, normalized event records, the qualitative codebook, and evidence for every reported case to support reproduction and reanalysis.
Agent execution is stochastic, and the substantial cost of long-horizon evaluation limits large-scale repeated runs; the authors note that Codex's sequential LoopsBench results are close to originally reported results and that the original LoopsBench study reports stable overall trends across repeated attempts. Manual analysis may introduce subjective judgments, which the authors constrain through a 50-trajectory pilot codebook, independent double annotation, recorded reasoning, and senior-author adjudication. On external validity, findings are scoped to the tested agents and benchmarks, and substituting another model could disrupt compatibility because agentic post-training may adapt a model to its scaffold's tool interfaces. In addition, no sampled ProgramBench task is fully resolved by any agent, so conclusions there rest on partial progress measured by test pass rate, and the direction of concurrency's effect differs across benchmarks and agents, leaving how task structure determines that direction an open question.
