Rejected patches stop going to waste: feeding verifier feedback into the next attempt lifts SWE-bench resolution from 38.9% to 41.7%
Lead
Repository-level bug-fixing agents can now be trained to treat a failed patch and its verifier feedback as context for the next attempt, raising first-attempt resolution on all 500 SWE-bench Verified tasks from GRPO's 38.9% to 41.7%.
Story
Training for repository-level software agents gains a recovery trajectory: after a patch fails verification, the repository returns to its original task state and the failed patch plus the verifier's diagnostic output become context for the next attempt. Reinforcement learning for these agents previously sampled several trajectories per issue independently and compared only their terminal rewards within a fixed group, so the failing tests and tracebacks from a rejected patch were never reused by a later attempt. On all 500 SWE-bench Verified tasks, with Qwen3.5-4B and SWE-agent under one feedback-enabled controller and attempt budget, FC-SWE reaches 41.7% Resolved@1 and 52.8% Resolved@2 against 38.9% and 48.5% for GRPO.
Each trajectory keeps only its own verifier outcome as its reward, while the comparison group pools the initial and recovery trajectories actually executed for the same issue. Broadcasting the final chain outcome to every trajectory would credit a rejected patch whenever a later recovery succeeded, and optimizing only the last trajectory would drop the rejected behavior that gives the successful recovery its contrast. Removing training-time failure context drops the same evaluation to 40.4% and 47.9% and conditional recovery from 19.0% to 12.6%, while a shared-reward variant reaches 17.7% conditional recovery against 20.4% for trajectory-local rewards.
What to watch
A next step is to attach recovery training to verifiers available at deployment, such as learned critics, agent-generated tests, or hybrid verification, and to separate the verifier that guides recovery from held-out tests used for final scoring. Evaluating such systems should measure verifier cost, false acceptance, false rejection, and diagnostic usefulness, because detecting a wrong patch and explaining how to repair it are distinct requirements.
The evaluation assumes benchmark-test feedback is available after every attempt, which need not hold in deployment. Whether recovery training still works when feedback is incomplete, noisy, or misleading, and whether the policy checks evidence rather than following it uncritically, remain open. The controlled training uses Qwen3.5-4B and one configuration, some ablations use a single decoding seed, and training compute is not matched to the baseline.
