Skip to main content
Back to timeline
arXivSource publication:

WideSWE tests coding agents on 120 cross-repository tasks: best configuration solves only 42.50%, and joint execution beats independent runs on requests

Synopsis

Drawing on 1,729,171 pull requests across 103 software ecosystems, the authors built WideSWE, 120 real tasks (60 bug fixes, 60 features) that require coordinated changes across multiple repositories, and evaluated seven coding-agent configurations: full task success ranges from 10.83% to 42.50%, with Codex CLI paired with GPT-5.6-sol highest, while joint and per-repository independent execution each help in different ways.

AI-generated editorial illustration: WideSWE: Can Coding Agents Coordinate Changes Across Repositories?

Interpretation

The paper introduces WideSWE, which formulates cross-repository coordination as an executable evaluation task: each task involves at least two scored target repositories, and a task counts as solved only when every target repository passes all its fail-to-pass and pass-to-pass tests. Earlier benchmarks such as SWE-bench, DeepSWE, and ProgramBench largely assess completion within a single codebase; WideSWE makes cross-repository coordination itself the object of evaluation. Tasks come from merged pull requests that explicitly reference another repository in the same ecosystem across 103 active ecosystems, retained after manual review and executable validation: 120 tasks covering 41 ecosystems, 253 target-repository instances, 2,815 F2P and 22,139 P2P tests.

Across seven agent configurations, full task success ranges from 10.83% to 42.50%, with Codex CLI paired with GPT-5.6-sol highest; that configuration solves at least one target repository in 83.33% of tasks but fully solves only 42.50%. The results separate repository-level progress from task-level completion, showing that passing in one repository does not mean the cross-repository request was fulfilled. Seven configurations are evaluated on 120 tasks with case-, repository-, F2P-, P2P-, and API-request-level metrics; among unresolved tasks, most (52.08%–69.57%, except Gemini) have at least one repository passing all F2P tests.

Failure trajectories fall into three categories: incomplete scope identification, recognized work without delivery, and post-edit failure, with post-edit failure most common for most configurations. Rather than reporting pass rates alone, the paper uses final diffs and trajectory evidence to locate where completion breaks down, with a two-stage annotation procedure reaching 91.2% agreement. For Codex CLI–GPT-5.6-sol the three shares are 37.68%, 2.90%, and 59.42%; for Gemini, recognized work without delivery reaches 52.34%, of which 82.14% remains in investigation, solution planning, or preparatory checks.

On 89 prompt-equivalent paired tasks, joint execution solves 40.45% versus 35.96% for independent execution; independent execution raises bug-fix success from 34.48% to 44.83% but lowers feature success from 43.33% to 31.67%, while using 296.6 versus 94.7 API requests per task on average. The paper directly compares one joint run in the ecosystem workspace against one separate run per target repository under identical prompts, and shows cross-repository context supports both implementation and verification. The paired comparison covers 89 tasks (29 bug fixes, 60 features); among repositories that fail in joint execution and are left unmodified, 60.00% succeed independently, versus only 9.43% of those already modified.

Perspective

The benchmark targets scenarios where one feature or fix requires coordinated changes across two or three repositories, evaluated in Linux environments, with workspaces that also expose up to 20 unscored context repositories (median three per task). It is suited to comparing agent configurations on cross-repository coordination and to studying differences between joint execution and per-repository independent execution. The paper states that results do not establish how agents perform on tasks involving more repositories or requiring other platforms.

The paper notes that the joint-versus-independent comparison matches per-run resource limits rather than the total budget available for each multi-repository task, so it examines whether solving repositories separately alleviates cross-repository difficulties rather than isolating execution scope under equal total budgets. The recovery-rate difference between repository groups is described as descriptive, not causal. Differences by task type and language mix (for example, multi-family bug fixes generally outperform single-family bug fixes while features do not follow the same pattern) come from this sample, and the paper cautions against generalizing them. In addition, task mining and selection involve multiple rounds of manual review, so the details of individual case choices require consulting the appendix.

Sources