POPO learns only from positive rollouts via online self-distillation, reaching 36.67% on AIME 2025 with Qwen-Math-7B versus GRPO's 30.00%
Synopsis
The work proposes Positive-Only Policy Optimization (POPO), an on-policy self-distillation framework integrated with RLVR that learns exclusively from online positive rollouts, using bounded importance sampling, a Siamese policy network with a momentum-based adaptation law, and a bounded similarity penalty instead of KL divergence; across MATH-500, AMC23, AIME 2024/2025, and Olympiad it outperforms GRPO, with Qwen-Math-7B reaching 36.67% on AIME 2025 versus GRPO's 30.00%.
Figure 1: (A) Scheme of RLVR for mathematical question-solving tasks. A general on-policy RL algorithm will reinforce the positive rollouts, while penalizing the negative rollouts. However, the negative rollouts may not contain a meaningful reasoning chain-of-thought in addressing the graduate-level mathematical questions, which motivates the POPO algorithm design. (B) A comparison of POPO and GRPO on the reward trace (upper panel) and mean output completion length (lower panel) on a simple GSM8K dataset.
arXivInterpretation
It proposes POPO, an on-policy self-distillation and RLVR framework that learns exclusively on online positive rollouts, with no disjoint negative rollouts used for gradient guidance during post-training. Compared with RLVR practice that relies on both positive and negative rollouts for reward signal, this work confines the learning signal entirely to the positive rollout set, noting that negative rollouts may admit no gradation of failure severity and that penalizing a few sampled negatives is unlikely to cover a meaningful reward signal under sparse binary rewards. The abstract describes the method: bounded importance sampling over the positive rollout set, with implicit negative gradients said to emerge naturally through reinforcing the positive probability via rollout redistribution; the derivation itself is not expanded in the provided text.
It stabilizes policy optimization through self-distillation: a Siamese policy network with a momentum-based adaptation law for asynchronous policy evolution, and a bounded similarity penalty term in the Siamese representation space replacing KL divergence. It replaces the usual KL-divergence constraint with a bounded similarity penalty in representation space and introduces a momentum-based adaptation law to support asynchronous policy evolution. The abstract explicitly lists both designs and states that ablation and sweep studies illustrate their necessity and robustness; the specific ablation settings and numbers are not given in the provided text.
Experiments on publicly available, well-established text LLMs across all-level mathematical benchmarks (MATH-500, AMC23, AIME 2024/2025, Olympiad) show POPO outperforming GRPO, with Qwen-Math-7B reaching 36.67% on AIME 2025 versus GRPO's 30.00%. It provides a direct comparison against GRPO and reports a concrete score gap on AIME 2025. The abstract reports benchmark names, the comparison method, and one concrete numeric comparison; full result tables, repetition counts, and statistical details are not present in the provided text.
Perspective
The work targets post-training of text LLMs with verifiable rewards, especially mathematical reasoning, with experiments on public benchmarks including MATH-500, AMC23, AIME 2024/2025, and Olympiad, comparing against GRPO using public models such as Qwen-Math-7B. It is relevant to researchers and practitioners who want to reduce reliance on negative rollouts or stabilize policy optimization via self-distillation. The setting presumes the existence of decidable positive rollouts and binary reward signals; whether it remains effective when rewards are not verifiable or positive samples are extremely scarce is not addressed in the text.
The provided text is the abstract and does not include full result tables, the specific settings and numbers of the ablation and sweep studies, training cost, repetition counts, or variance information, so the stability of the 36.67% versus 30.00% gap cannot be judged. How implicit negative gradients arise from rollout redistribution, along with its derivation and conditions of applicability, still requires the main text. The form of the bounded similarity penalty, its bound values, and the hyperparameter sensitivity of the momentum adaptation law are open questions for further reading. In addition, how positive-only learning performs on tasks with scarce positive samples or noisy rewards is not covered in the text.
