Skip to main content
Back to timeline
arXivSource publication:

GRAFT swaps trajectories between two heterogeneous models, beating GRPO for both at equal budget with a 2.1-point average gain

Synopsis

The work proposes GRAFT, a framework that in Reinforcement Learning with Verifiable Rewards (RLVR) replaces a receiver's all-fail rollout groups with peer groups containing both successful and unsuccessful responses, controlling cross-model mismatch through sequence-level compatibility weighting and token-level importance ratio clipping; across three heterogeneous model pairs and five mathematical reasoning benchmarks, GRAFT improves both models over GRPO at the same per-model rollout budget, gaining 2.1 points on average and up to 4.5 points, with 1.8 points on average preserved when reusing stored peer trajectories.

AI-generated editorial illustration: Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR

Interpretation

GRAFT reframes the RLVR problem of all-fail groups carrying no gradient signal as a cross-model complementarity problem: when a receiver's entire rollout group fails on a prompt while the peer produces both successful and unsuccessful responses on that prompt, the receiver's failed group is replaced by the peer's full group, retaining the peer's own source-computed group-relative advantages. Prior GRPO-style methods depend on the learner's own exploration: dynamic sampling discards zero-variance groups, larger groups raise success odds at higher cost, and advantage reshaping can only suppress sampled failures. HACPO shares peer rollouts more broadly, and SGT transfers only verified peer successes through a fixed-weight supervised loss. GRAFT treats which prompts to select and how to learn from them as two interacting decisions addressed together. The paper reports complementarity in independent GRPO runs of SmolLM3-3B-Base and Qwen3-1.7B-Base: SmolLM3-3B-Base solves 47.9% of the prompts on which Qwen3-1.7B-Base fails across all eight rollouts, and the reverse direction is 18.7%. The main experiments cover three heterogeneous model pairs and five mathematical benchmarks, with each checkpoint evaluated over five runs and eight samples per prompt, reporting mean and standard deviation.

GRAFT controls cross-model mismatch with two layers: a sequence-level compatibility gate that assigns a bounded weight from average token log-likelihoods with a floor threshold, and token-level importance ratio clipping that tracks only the receiver's own change during optimization, with peer responses re-tokenized under the receiver's tokenizer and entering the same objective. The paper factorizes the string-level likelihood ratio into the receiver's change during optimization and its initial mismatch with the peer, handling the former with a token-level PPO surrogate and the latter with a sequence-level compatibility weight; the authors state explicitly that the compatibility score is an empirical proxy from average token log-likelihoods, not an exact cross-tokenizer density ratio. Ablations show that removing the compatibility gate lowers SmolLM3-3B-Base by 8.36 points and Qwen3-1.7B-Base by 2.51 points on average; removing only the floor lowers Qwen3 by 3.56 points; replacing the token-level ratio with a sequence-level ratio also reduces performance. In the threshold sweep, 0.5 achieves the highest aggregate score for both receivers.

The processing order of peer minibatches is itself a control mechanism: placing minibatches containing peer groups after the receiver's own on-policy minibatches makes token-level clipping actually activate on peer tokens. The paper notes that before the first optimization step on a newly collected rollout batch the importance ratio is one, so a grafted minibatch processed first would enter unclipped and the sequence-level weight would be the only control on mismatch; peer-last ordering gives the receiver's own data priority and turns clipping into a second mechanism for moderating peer influence. Appendix H measures the clipped-token fraction over training steps 1-48: peer-last yields 0.28%-0.44% on grafted responses, uniform 0.13%-0.25%, and peer-first below 0.01%; peer-first has zero clipping on peer tokens in 97.6% of steps versus 0.0% for peer-last. In ablations, peer-last scores higher on average than peer-first and uniform.

The gains do not strictly require synchronous co-training: using peer responses and generation log-probabilities stored from the partner's independent GRPO run, with rewards re-verified and the same selection, weighting, and update rule applied, improves all six model blocks over GRPO by 1.8 points on average, about 84% of the 2.1-point online gain. Online co-training keeps both models and their optimizer states resident; the stored-trajectory version shows most of the benefit of cross-model exploration persists without loading or training the peer model, although transfer becomes unidirectional and balanced exchange is inactive. Stored-trajectory experiments improve all six model blocks by 0.80-2.68 points; on Pair 1, obtaining both selected receiver checkpoints requires 23.2 GPU-hours excluding the prior GRPO runs used to collect peer logs, versus 40.9 GPU-hours for online GRAFT. Across three independent training runs per pair, GRAFT maintains higher mean aggregate performance than GRPO in all six model blocks.

Perspective

The result targets research and engineering settings that use verifiable rewards and GRPO-style post-training, especially teams holding several heterogeneous open-source base models that want to complement each other's exploration without designating a stronger teacher. The method applies to two models trained as a pair: replacement triggers when the receiver's group fails entirely and the peer group contains both successes and failures, and the compatibility threshold selected on Pair 1 is fixed for the other two pairs. For practitioners seeking lower memory and scheduling cost, the stored-trajectory version shows that responses and log-probabilities left by a partner's independent GRPO run can be reused without keeping both models resident. The paper attributes the gains to complementarity itself, so pairs with stronger complementarity gain more.

The scope limits the paper itself lists are worth watching: the size of the gain depends on how complementary the two models are and is smallest on Pair 3; across tokenizers the compatibility score is a proxy rather than a density ratio; the study covers only two-model pairs, only math, and only base models of at most 3B parameters. Exchange among more than two peers and domains without verifiable rewards are left to future work. In addition, main-table results use the first training run, and while the three repeats in Appendix C show GRAFT maintaining higher mean aggregate performance than GRPO in all six model blocks, the spread across individual runs is worth observing during replication. This evidence bundle is the full text, including main tables, ablation tables, compute accounting, and appendix statistics, so the judgments above rest on the complete content the paper reports.

Sources