Skip to main content
Back to timeline
arXivSource publication:

Xiaomi's groupwise agentic grading and advantage redistribution lifts 310B and 1.02T code agents to 67.9 and 71.9 on DeepSWE

Synopsis

Xiaomi introduces GAGAR, a framework for code agent RL in which an SFT-trained agentic grader jointly inspects test-passing trajectories within a task group, ranks them along five quality dimensions, and redistributes advantage from lower- to higher-quality solutions in a sum-preserving way; in controlled code-only experiments with MiMo-V2.6-Flash (310B total parameters) it reaches a 62.2% DeepSWE v1.1 pass rate at step 28, 12.1 percentage points above the binary-reward baseline, and after industrial-scale mixed-task RL with Flash and Pro (1.02T total parameters) the checkpoints reach DeepSWE v1.1 avg@3 of 67.9 and 71.9.

AI-generated editorial illustration: Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Interpretation

GAGAR replaces independent per-trajectory scoring with joint within-group comparison: the grader sees the task specification, repository, full trajectories, submitted patches, and test outputs in a shared workspace, first reviews a turn-by-turn summary to locate parts needing closer inspection, then cross-checks patches against repository code and test logs, can run targeted checks, and must cite patch locations, trajectory events, or execution results for negative assessments. Prior reward modeling and grading work often relies on fixed summaries or single-pass text assessment, such as the Groupwise Ranking Reward cited in the paper, which uses a single-pass, text-based grader; GAGAR instead performs interactive inspection of complete implementations for the same task and uses failed trajectories as context that reveals failure modes. The paper provides the grading procedure, five criteria (approach suitability, implementation precision, minimality of changes, unintended side effects, codebase consistency), and three quality tiers with within-tier ranking rules; the grader is SFT-trained, and the paper reports an average end-to-end grading time of about 600 s versus about 2,000 s for the initial Claude Opus 5 implementation.

GAGAR introduces sum-preserving advantage redistribution: lower-ranked passing trajectories are first downweighted by quality factors, then all passing trajectories are rescaled by a common factor so that the total positive advantage of the passing subset is restored, failed-trajectory advantages are unchanged, and the quality ordering and relative weights are preserved. The paper notes that downweighting alone removes positive credit without changing negative advantages, making total negative advantage exceed total positive advantage; in the ablation, the downweighting-only run averages a policy-gradient loss of 0.0304 over steps 1–30 versus 0.0021 for the full method, with entropy rising from 0.359 to 0.905 and mean training-rollout length from 47.1k to 114.1k tokens, compared with 0.358 to 0.513 and 46.5k to 69.9k tokens for the full method. This comes from a controlled ablation in the same Flash setting with concrete loss, entropy, and length values; the paper also gives both an advantage-space and an equivalent reward-space formulation and states that exact conservation does not hold when the rescaling cap is active.

In controlled code-only RL, GAGAR improves task performance and interaction efficiency relative to the binary-reward baseline: the baseline is stopped at step 28 due to rapid degradation (DeepSWE pass rate falling from 58.5% at step 20 to 50.2%), while GAGAR reaches 62.2% at step 28, 12.1 percentage points higher, and peaks at 63.4% at step 44; on SWE-bench Pro the baseline plateaus at about 59% from step 16 to 28, while GAGAR reaches 62.5% at step 52. The paper attributes the gains to both the quality signal and training stability, and reports fewer turns and tokens at shared checkpoints: at step 28 DeepSWE mean turns fall from 132.3 to 111.6 and token length from 191.9k to 172.9k (reductions of 15.6% and 9.9%), while SWE-bench Pro turns fall from 58.4 to 54.6 and tokens from 79.9k to 68.6k (reductions of 6.5% and 14.1%). Results come from a controlled code-only Flash comparison with training batch size 128 and 16 rollouts per prompt, evaluating three samples per task and reporting avg@3; the paper states that baseline and GAGAR are compared under identical settings and at shared training steps.

In industrial-scale mixed-task RL, GAGAR is integrated into training of MiMo-V2.6-Flash and MiMo-V2.6-Pro, whose final checkpoints reach DeepSWE v1.1 avg@3 of 67.9 and 71.9 and SWE-bench Pro scores of 60.9 and 62.7; the paper reports that Pro, with 1.02T total parameters, outperforms Kimi K3 at 2.8T (DeepSWE 69.0) and surpasses GPT-5.6 Sol on SWE-bench Pro with 62.7% versus 60.5%, while approaching GPT-5.6 Sol's 73.0% and Claude Opus 5's 74.0% on DeepSWE. The paper emphasizes applying quality-aware redistribution to coding-task groups within a heterogeneous workload rather than only to code-only training, and states that the MiMo-V2.6 series thereby achieves code capabilities competitive with frontier models. This part reports industrial-scale runs using 1,568 prompts per update with 16 rollouts per prompt, with grading and redistribution operating on eligible coding-task groups; the comparison numbers come from Table 1 of the paper.

Perspective

This work targets code agent RL settings where executable tests provide rewards and trajectories can span hundreds of interaction turns, and it fits training pipelines that have repository context and execution environments; the paper validates on pre-RL SFT checkpoints of MiMo-V2.6-Flash (310B total, 15B active parameters) and MiMo-V2.6-Pro (1.02T total, 42B active parameters), integrating grading and redistribution into eligible coding-task groups within mixed-task RL. For teams aiming to reduce developer review and rework and to produce precise, merge-ready implementations, it offers a path to spend additional computation on graders that identify which behaviors are worth learning; the paper also mentions ongoing exploration of turn-level credit assignment and applying zero-sum redistribution to different graders, penalty mechanisms, and advantage adjustment strategies.

The grader is SFT-trained, and the paper reports an average grading time of about 600 s with good accuracy, but how grading quality varies across task distributions and how often unusable grading results fall back to original advantages still need observation in more settings; the paper states that exact conservation does not hold when the rescaling cap is active, in which case the total positive advantage and failed-trajectory advantages need not be preserved, and the practical impact of that branch is worth watching; the quality audit is based on a fixed random sample of 30 DeepSWE tasks reviewed jointly by Claude Opus 5, a limited sample size; the mixed-task RL comparison numbers come from Table 1, and training and evaluation details of the external baseline models are outside this paper's scope. In addition, this reading is full text, but formulas and tables appear as placeholders in the text, so specific coefficients and complete table values should be checked against the original.

Sources