RucTangle splits coding-agent patches into commit histories that stay runnable at every commit and lifts pass@1 by 16.6% on average in regression repair
Synopsis
The work introduces RucTangle, an agentic commit-untangling method that groups a patch's changes by development purpose, orders them, applies each commit and checks whether the code still runs, revising order first, then regrouping via the agent, and finally merging adjacent commits while preserving the patch's final code; on 131 agent-generated patches it is the only method that reconstructs every patch with all commits runnable, at a median of eight commits, whereas baselines leave 20.6%–37.4% of histories with at least one unrunnable version. The authors also present TangleEval, which has two coding agents repair 453 regression-introducing patches; adding RucTangle histories plus git bisect results raises pass@1 from 26.9% to 31.3% and from 30.8% to 36.0%.
Figure 1 : A simplified patch with two development purposes: adjusting the logging level (A) and switching the default backend from local to sqlite (B and C). Change B replaces the entry in the backend dictionary, and Change C updates the default selection. Because bootstrap.py evaluates BACKENDS[DEFAULT_BACKEND] at import time, separating B and C causes a KeyError between these commits. Keeping A separate and grouping B with C avoids this mismatch.
arXivInterpretation
RucTangle reconstructs all 131 agent-generated patches and keeps the code runnable after every commit (pytest can discover and load tests), with a median of eight commits and only 0.8% single-commit histories. Prior methods (FileSplit, HunkSplit, Armchr, Atomizer) reconstruct 96.2%–100% of patches but only 62.6%–79.4% of histories are runnable after every commit; the paper makes runnability an explicit constraint and uses execution feedback to revise grouping and order. Comparison against four baselines on 131 Claude Opus 4.6 patches across 100 Python repositories in SWE-CI; runnability is defined as pytest discovering and loading tests, checked per version in separate containers.
RucTangle histories temporarily break fewer tests: among tests passing in the final version, an average of 1.7% transition from passing to failing after some commit, versus 16.4%–32.4% for the four baselines; test gains are also less concentrated (MTP 0.552 vs 0.610–0.739, TPE 0.562 vs 0.361–0.419). The paper uses the distribution of test gains as a proxy for development purpose, extending evaluation from syntactic agreement with reference groups to functional decomposition and temporary regressions along the commit sequence. Statistics on the same 131 patches; RucTangle does not use test pass/fail results when revising grouping or order, so this is not a directly optimized objective.
On 453 regression-introducing agent patches, adding RucTangle histories plus the git bisect candidate commit to the repair agent's context raises pass@1 from 26.9% to 31.3% for Qwen3-Coder-Next and from 30.8% to 36.0% for Nemotron, a 16.6% average relative improvement; Armchr and Atomizer yield smaller average relative gains of 4.5% and 6.1%. Prior evaluation of commit untangling compared generated groups with developers' original commits syntactically and did not directly show maintenance benefits; TangleEval makes whether an untangled history helps a coding agent repair regressions a measurable outcome. Two models on 453 patches with three attempts per patch per condition, using the same failing code, tests and repair budget across conditions; four additional models were evaluated on a shared 150-patch sample with one attempt each, where two improved, one was unchanged and one slightly decreased.
Trajectory cases show agents use commit diffs and earlier versions to decide what to repair and consult commit messages to decide which behavior to preserve; one generated message calling deleted newline handling "redundant" led an agent to repeatedly question a correct repair and exhaust its step budget. The paper links commit-message quality to subsequent agent decisions, noting that misleading messages can undercut the value of untangled histories, and proposes designing project histories for subsequent agents. Case studies of trajectories from the regression-repair evaluation, selecting patches where NoHistory and RucTangle produced different outcomes for the same model and attempt index.
Perspective
The result targets teams maintaining Python projects with coding agents: splitting a large agent-produced patch into purpose-grouped commits that stay runnable after each commit, and supplying the git bisect candidate commit and commit messages to the repair agent, improves regression repair without retraining models. Runnability is defined as pytest discovering and loading tests, so it applies to Python projects with collectable tests; the paper lists the effectiveness of this checking stage for other languages or deployment requirements as open. The method preserves the patch's final code, does not require functional tests to pass, and does not repair errors introduced by the original patch, which is a separate task.
Repair benefit varies by model: four of six improve, one is unchanged and one slightly decreases, so a single improvement figure should not be read as universal. Most models also introduce new test failures more often while fixing the target regressions, so the net pass@1 gain is smaller than the rate at which target failures are fixed. The merged-history experiment evaluates only Nemotron, and merging changes both the scope of the bisect candidate commit and the messages presented together, so the result reflects the combined change in available information rather than commit count alone. Message quality was not independently controlled, leaving its effect an open question. The repair comparison evaluates the complete RucTangle system including bisect localization, without isolating component contributions.
