FastRL brings about 2.07x training speedup and roughly 1.64% average accuracy gain to GRPO-style methods via advantage-aware pruning and adaptive sampling
Synopsis
The work proposes FastRL, a plug-and-play reinforcement learning framework that selectively preserves high-advantage trajectories while maximizing inter-trajectory gradient diversity through an advantage-aware pruning strategy, and dynamically adjusts sampling scale across training stages based on historical pruning distributions via an adaptive rollout sampling mechanism, achieving an average 2.07x training speedup on Geometry3K and GeoQA8K-R1V along with approximately 1.64% improvement in average accuracy on visual reasoning benchmarks, and integrating seamlessly into GRPO, DAPO, and GSPO variants.
Figure 3: Training efficiency and parameter sensitivity of FastRL on Geometry3K (top) and GeoQA8K-R1V (bottom). (a1)-(a2) show the the rollout sampling scale G t G_{t} ; (b1)-(b2) present the number of trajectories used for policy updates; (c1)-(c2) show the sensitivity to the hyperparameter τ \tau .
arXivInterpretation
An advantage-aware pruning strategy is introduced to selectively preserve high-advantage trajectories while maximizing inter-trajectory gradient diversity. Addressing the degradation of downstream learning signal efficiency caused by low-information or highly homogeneous trajectories, the strategy preserves high-advantage trajectories while also accounting for gradient diversity, rather than filtering by a single advantage criterion. The abstract describes the design goal and mechanism of the strategy; specific pruning criteria and diversity measures are not given in the abstract.
An adaptive rollout sampling mechanism is designed to dynamically adjust the sampling scale across different training stages based on historical pruning distributions. Addressing the computational overhead of per-question multi-rollout sampling, the mechanism uses historical pruning distributions as a signal to adjust sampling scale, balancing exploration adequacy and computational efficiency. The abstract states the basis and goal of the mechanism; specific adjustment rules and thresholds are not given in the abstract.
FastRL integrates seamlessly into GRPO, DAPO, and GSPO variants, achieving an average 2.07x training speedup on Geometry3K and GeoQA8K-R1V and approximately 1.64% improvement in average accuracy on visual reasoning benchmarks. Relative to original GRPO and its variants, the framework improves both training efficiency and policy learning effectiveness rather than optimizing only one of them. The abstract reports average speedup and average accuracy improvement values across multiple variants and datasets; specific experimental configurations and statistical details are not expanded in the abstract.
Perspective
The work targets researchers and engineering practitioners who train policies with group relative policy optimization methods such as GRPO, DAPO, and GSPO, and applies to settings that require per-question multi-rollout sampling and care about training overhead and learning signal efficiency. The results described in the abstract are obtained on Geometry3K and GeoQA8K-R1V and on visual reasoning benchmarks, and integration with three GRPO variants is verified; its plug-and-play positioning means it can be substituted into or layered onto existing training pipelines.
The abstract does not give the specific criteria of advantage-aware pruning, the measure of gradient diversity, the adjustment rules and thresholds for adaptive sampling scale, nor the model scale, training steps, and statistical significance of the experiments. Readers concerned with these implementation details and result robustness need to consult the full text; in addition, the abstract mentions that source codes will be released, and actual availability remains to be confirmed.
