Skip to main content
Back to timeline
arXivSource publication:

Treating parallel coding-agent coordination as scheduling: NP-Bench lifts clean integration from 1/9 to 9/9 and rescues a live breaking contract change from 0.00 to 1.00

Synopsis

The authors reframe coordination among parallel coding agents as a scheduling problem: before agents start, partition work into disjoint scopes from each item's declared scope and order merges along the producer-consumer dependency graph; they implement this planner in Nerveplane and evaluate it with NP-Bench, a three-arm benchmark grounded in a real git merge, where deterministic clean-integration success rises from 1/9 to 9/9 and merge conflicts fall from 13 to 0, and a live breaking contract change rises from 0.00 to 1.00 on a frontier model and 0.60 on a small one.

Source-provided article image: Verifying Coordination in Parallel Coding Agents: NP-Bench and a Scheduling Planner
Figure 1 ·

Figure 1: Nerveplane observes each agent’s worktree by polling git (passive sensing), then routes only the relevant events and, with the planner, hands each agent a disjoint scope and a merge slot before the edits happen, rather than warning after a collision is sensed.

arXiv · Page 4

Interpretation

NP-Bench is an environment-grounded three-arm verifier (no coordination C0, reactive detection C1-detect, proactive planning C1-plan) that reads ground-truth signals off a real git merge, including clean-tree success rate, merge conflicts, wasted LOC, scope leakage, consumer-adaptation rate, and repeated-mistake rate. Existing coding-agent evaluation such as SWE-bench, HumanEval, and LiveCodeBench scores single-agent, single-repository task success and does not measure whether parallel agents integrate cleanly; NP-Bench adds that coordination axis and turns 'did coordination help' from an assertion into a measurement. Tier A is a deterministic simulation driving the real registration, worktree sensing, conflict detection, and planner core over scripted scenarios, scored by a real git integration, CI-gated and reproducible; Tier B runs real coding agents headless in isolated git worktrees, integrates their branches with a real merge, and is nondeterministic with Wilson 95% intervals reported.

Across nine scripted scenarios the planner raises clean-tree success rate from 1/9 under C0 and 4/9 under C1-detect to 9/9, drives total merge conflicts from 13 to 0 and wasted LOC from 54 to 0; under concurrent contention the reactive detector collapses to the no-coordination baseline while the planner holds at zero, and the advantage grows monotonically with the number of agents. Reactive coordination ties the planner on sequential dependency scenarios, where the warning has time to land, but under agent-speed concurrency the warning always arrives after the edit; the planner prevents the collision by construction through up-front disjoint partitioning and topological merge order. The deterministic three-arm suite covers shared-file, contract, cross-repo fan-out, an independent control, and concurrent/high-contention variants; cross-repo routing is exact on the fan-out scenario, warning the two real consumers and not the unrelated service (100% routing hit, 0% false positives); on the control scenario all three arms succeed and no warnings are raised.

In the live contract-migration scenario, where a producer renames an invoice money field from dollars amount to integer cents amountCents and a consumer that fails to adapt merges cleanly but is semantically broken, the planner lifts clean-integration rate from 0.00 to 1.00 on frontier models and to 0.60 on a small model, and to 1.00 on a second vendor's model, while C1-detect sits at 0.00 on every model. A vague just-in-time warning ('a teammate may be editing') is worth about as much as nothing to a live agent: raw replies show the consumer knew something may change and still coded to the stale field; only the explicit up-front assignment carries the contract shape the agent needs. Live runs span two model families (Claude Opus 4.8, Claude Haiku 4.5, and OpenAI GPT-5-Codex), with the pivotal contract cell run at a larger sample where the planner's Wilson interval separates cleanly from the baselines'; the C0 failure is not a capability failure because the needed information is simply absent from the consumer's worktree.

On the shared-file scenario, scope leakage for reassigned frontier-model agents is 0.00, meaning every planner agent honored its assigned scope; in a separate cross-session memory experiment the repeated-mistake rate falls from 1.00 to 0.00 on both a strong and a weak model. This answers whether 'disjoint scopes don't conflict' is a tautology: real agents do respect the scopes they are handed; and the memory effect is an information effect, because the sanctioned choice is not recoverable from the repository and model strength cannot substitute for it. Scope leakage is measured at 0.00 on the shared-file scenario with frontier agents, several of which noted they were acting 'as coordinated'; the memory experiment uses a sandbox repository with two symmetric authentication backends whose sanctioned one appears nowhere in the code, with session 1 recording the decision and a fresh session 2 asked which backend to import.

Perspective

The result is aimed at teams and platform builders running several coding agents in parallel on the same codebase, in settings where each work item's file scope and consumed contracts can be declared and the dependency relation is a DAG. In that setting the planner turns conflict cleanup into conflict prevention: concurrent work receives disjoint scopes and producers merge before their consumers. The authors note the benefit comes from how work is allocated rather than from model reasoning, so across the two capability tiers and two vendors tested it did not shrink as models got stronger. The system and the NP-Bench harness are open source (Nerveplane v0.17.0); Tier A is deterministic and CI-gated, serving as a regression check on mechanism and routing precision, while Tier B supplies live-agent behavior evidence. The cross-session memory result indicates that recording a team decision in a queryable ledger stops an agent in a later session from repeating the same mistake.

Scope inference remains unbuilt: the declared scopes in these experiments are hand-seeded in the harness, the planner's guarantee is only as good as those scopes, and inferring them automatically from a symbol graph or the service graph is future work. Scope leakage is measured only on the simple shared-file case, and leakage could be higher on subtler scopes. The C1-detect warning is deliberately vague, which the authors argue is realistic rather than a straw man, but a detailed-but-late ablation separating specificity from timing is missing. Most live cells besides the pivotal contract cell are pilots with wider intervals and the scenarios are hand-constructed; larger samples across every cell and a broader scenario set are the obvious next runs. Separately, routing facts to agents did not improve long-context retrieval accuracy at window-fitting scales, its value being cost and capacity, so stronger claims for context routing would need additional evidence.

Sources