Splitting repository-level SWE tasks into category experts and distilling them back into one model lifts Pro-618 mean resolution to 58.04%
Synopsis
Addressing a "category see-saw" in repository-level software engineering, where pooled agentic reinforcement learning improves some task categories while regressing others and aggregate resolution hides the change, the work builds a category-aware expert-training and policy-integration framework: SWE Labeler groups tasks by repository domain into service/data/security (A), user-facing applications (B), and systems/tooling/runtimes (C); Agentic-miniRL with a Refresh–Repair–Expand loop trains same-origin category experts; and label-routed multi-teacher on-policy distillation (MOPD) consolidates them into one deployable student, which reaches 58.04% mean resolution on Pro-618 (+5.39 points over base) and 59.00% on SWE-bench Multilingual (+2.
Interpretation
The paper partitions repository-level SWE tasks into categories using observable labels and shows that pooled joint RL conceals opposing movements across those categories. Prior cross-domain expert splitting and fusion relies on domain boundaries that are separable by construction, such as mathematics or code generation; here, within a single software engineering domain, category granularity is established by an evidence-grounded hierarchical labeling system (26 Task Type L1 and 119 L2 labels, 21 Repository Domain L1 and 108 L2 labels, plus three four-level scale axes), which defines the "category see-saw" and a see-saw gap metric. On Pro-618, an audited 618-instance subset of SWE-bench Pro with 221/201/196 instances in A/B/C, Pooled RL shows gains in some categories coinciding with regressions in others, and a positive overall gain can coexist with a negative minimum category gain; Balanced RL reaches close overall performance with only 1,548 distinct training instances (516 per category) versus 6,723 for Pooled RL (55.34% versus 55.50%), with a higher average minimum category gain (0.19 versus 0.08 points) but no reduction in the average see-saw gap (1.18 versus 1.05 points).
The Refresh–Repair–Expand (RRE) loop lets category experts repeatedly consolidate their own successful trajectories, addressing average progress and instance-level regressions together. Unlike the common pipeline of synthesizing trajectories with a larger teacher, distilling by SFT, and then applying RL, all three experts start from the same base policy and acquire category behavior directly through executable Agentic RL, with no external model supplying solution trajectories or action targets; RRE couples mastery refresh, Repair SFT on verifier-approved successes, and expansion of the task frontier so that data selection co-evolves with the policy. Across 2,769 matched training records, initial RL raises mean success rates by 7.26, 6.76, and 4.60 points for A/B/C, yet 839 records (30.3%) decline; after Repair SFT the category means rise to 56.28%, 59.59%, and 52.75%, gains of 10.23, 14.05, and 9.47 points over initial RL, and of the 751 regressed records included in repair, 125 (69.8%), 103 (73.0%), and 264 (61.3%) recover to or above their base rates. On the corresponding Pro-618 categories, the first repair adds 3.77, 1.50, and 2.89 points over initial RL, and the final experts exceed base by 7.84, 4.48, and 5.61 points.
Label-routed multi-teacher on-policy distillation (MOPD) consolidates the three experts into one student that beats both joint-RL baselines overall and in every category. MOPD augments on-policy distillation with a ReLU-gated reward-extrapolation term that keeps only each teacher's improving direction over the reference; teachers are same-origin (shared tokenization, prompt format, base checkpoint, and agent interface) and routing is defined by observable category labels rather than given domain boundaries. The final MOPD policy reaches 58.04% mean Full resolution on Pro-618 across three evaluation rounds, +5.39 points over base, with Pro-A/B/C at 58.07%, 59.37%, and 56.63% (+6.33, +4.98, +4.76 points); versus Pooled RL it improves Full by 2.54 points and A/B/C by 3.32, 2.65, and 1.53 points, and versus Balanced RL by 2.70 points overall and 3.62, 2.32, and 2.04 points by category, giving minimum category lifts of 1.53 and 2.04 points. On SWE-bench Multilingual it reaches 59.00%, +2.78 points over base, with paired instance-bootstrap 95% intervals for Full differences of [0.33, 5.22] versus base and [0.67, 6.22] versus Pooled RL.
The integrated gains are uneven: the fraction of each expert's improvement retained differs markedly by category. The paper evaluates integration from two perspectives, gains over the joint-RL baselines and expert-gain recovery, noting that the category with the smallest final improvement need not be the one with the largest proportional integration loss. By gain recovery, MOPD retains 80.8%, 111.1%, and 84.8% of the corresponding expert gains on A/B/C; recovery above 100% on B means the student exceeds its expert by 0.50 points, while on A and C the student remains 1.51 and 0.85 points below the corresponding experts yet still above both joint-RL baselines.
Perspective
The framework targets post-training of repository-level software engineering agents: it applies to task pools with containerized repositories, issue-style problem statements, and executable fail-to-pass and pass-to-pass tests, and it needs an auditable category labeling scheme to drive training-pool construction, teacher routing, and category-level evaluation together. It lets follow-up work test the hypothesis of "split experts, then integrate" within a single domain, and it supplies directly reusable measures such as the see-saw gap, minimum category gain, and expert-gain recovery. For practitioners, this means that when aggregate scores look stable but category performance cancels out, organizing training streams by category is worth trying rather than only adjusting mixture proportions; the finding that Balanced RL reaches close overall performance with 1,548 tasks versus 6,723 for Pooled RL also suggests a smaller category-balanced training set can stay competitive in this setting.
The paper itself flags several open questions: tasks can span multiple categories while the current hard routing assigns each task to one expert category; training multiple experts and running teacher forwards adds substantial compute, and the best configuration may vary with data, base model, and budget; same-origin teachers may be too similar to add signal or, after heavy specialization, too distant for stable integration; benchmark pass rates do not fully measure patch quality, and test-driven rewards can invite evaluator hacking or reflect brittle task specifications. The authors also note that repository sanitization and fresh-sandbox replay prevent local-history leakage and rollout-state tampering, and trajectory audits screen external retrieval, but these controls do not eliminate pretraining memorization or every weak-verifier failure; results on Pro-618 cannot be generalized to an unexecuted full Pro set; the experiments evaluate the complete training framework and its stage-wise outcomes without isolating the contributions of each optimization component, replay choice, or routing granularity; and single-run category trajectories include evaluation variability, leaving broader comparisons of sampling and routing strategies open. In addition, some table values in the loaded text appear as placeholders, so the stage-wise Pro-618 scores after each Repair SFT phase and parts of the training-instance decomposition can only be read through the percentage-point changes described in the prose.
