EasyPPO traces two critic failure modes in PPO and stays stable across three tasks, gaining 14.89%, 2.28% and 9.47% over PPO
Synopsis
The work identifies two critic failure modes that destabilize PPO for LLM reinforcement learning — filtering truncated rollouts from both actor and critic shifts the policy objective to reward conditioned on completion, and heterogeneous return noise across prompts lets high-variance prompts dominate critic updates in finite batches — and introduces EasyPPO, which combines actor-only overlong filtering, noise-normalized critic regression weighting each prompt by the inverse standard deviation of its sampled returns, and moderately smaller critic mini-batches with gradient clipping, remaining stable across continuous-reward coding on FrontierCS, binary-reward mathematical reasoning on AIME24 and multi-turn search on Search-R1, with best validation scores improving over PPO by 14.89%, 2.
Interpretation
The paper shows that extending overlong filtering from the actor to the critic changes what the critic learns: instead of predicting the mean return over all rollouts, it predicts the mean return among completed responses, which in the accurate-critic limit shifts the policy objective to reward conditioned on completion, so truncation can rise while conditional reward still improves. Overlong filtering had mainly been used in critic-free methods such as GRPO to reduce reward noise from truncated rollouts, and systems such as Nemotron-Cascade exclude overlong rollouts in math RL; this work places that practice inside PPO's actor-critic framework and analyzes how filtering both sides changes the policy objective itself. The paper provides an idealized-limit derivation and compares three filtering strategies on FrontierCS: without filtering PPO shows repeated reward collapses; with joint actor-and-critic filtering reward among non-truncated rollouts rises but the truncation ratio approaches its limit and overall reward stays low; actor-only filtering is the most stable and achieves the highest overall score. In the AIME setting joint filtering also shows rising truncation and deteriorating overall reward.
The paper identifies a second failure mode in heterogeneous return noise within a batch: when the critic minimizes squared error, return noise leaves the optimal prediction unchanged but enters the expected squared gradient magnitude alongside prediction error, so as predictions improve the noise can dominate and prompts with more variable returns disproportionately influence finite-batch updates. Variance-weighted regression has been studied in RL and supervised learning, but this work uses the grouped sampling already present in LLM RL to estimate prompt-level return variance directly, without an auxiliary variance model, and applies the weighting to the critic loss rather than to actor advantages. The paper gives a gradient second-moment decomposition and proposes weighting each prompt's critic loss by the inverse standard deviation of its sampled returns, normalized to mean one over the prompt batch, with a variance floor for discrete rewards. In offline diagnostics on FrontierCS and AIME, prompts with more variable returns tend to contribute larger gradients without normalization, and this pattern is stronger at later checkpoints; normalization largely removes this dependence, though the authors note gradient outliers remain and the AIME trends are less uniform.
The paper introduces EasyPPO with three changes — actor-only overlong filtering, noise-normalized critic regression, and moderately smaller critic mini-batches with gradient clipping — and reports stable training and better performance than baselines across three tasks. Unlike VAPO's value-learning and advantage-estimation improvements or HL-Gauss PPO's categorical critic objective, EasyPPO keeps the standard PPO actor update and targets the critic's exposure to prompt-level return noise and the effect of filtering on the objective. Across 200–300 training updates on FrontierCS, AIME24 and Search-R1, EasyPPO is the only compared method stable on all three tasks, with best validation scores improving over PPO by 14.89%, 2.28% and 9.47%; on FrontierCS none of three EasyPPO seed runs collapses while two of three actor-only-filtering runs do. Ablations show noise normalization stabilizes training under both mini-batch configurations, whereas without it splitting the same rollout batch into four smaller mini-batches produces a sustained collapse.
The paper characterizes a trade-off in critic mini-batch size: smaller mini-batches confine an outlier's post-clipping influence to fewer rollouts, but each clipping operation faces larger stochastic fluctuations, motivating a moderate size. This analysis treats clipping granularity and noise averaging as competing effects rather than simply favoring larger critic batches. The paper derives an idealized fixed-parameter bound on a single outlier's influence and a per-mini-batch gradient variance relation, and varies the number of critic mini-batches per rollout batch on FrontierCS: the first-crossing step decreases approximately linearly with the number of mini-batches, while larger mini-batches achieve higher explained variance later in training.
Perspective
The result targets research and engineering settings that use PPO-style actor-critic methods for LLM reinforcement-learning post-training, especially tasks with continuous rewards and large differences in return variance across prompts; the experiments cover continuous-reward coding on FrontierCS, binary-reward mathematical reasoning on AIME24 and multi-turn search on Search-R1, with Qwen3.5-9B and Qwen3.5-9B-Base and training horizons of 200–300 updates. For a reader, the directly reusable content is three operational rules: apply overlong filtering only to the actor, weight critic regression by the inverse standard deviation of prompt returns, and choose a moderate critic mini-batch size while adjusting the learning rate accordingly. The paper also notes that when fewer rollouts per prompt are desired, offline profiling with online updates could provide variance estimates, which is left to future work.
The offline gradient diagnostics characterize gradients of the terminal-return MSE and do not replay training-time GAE targets or optimizer updates, and the authors state the fits summarize associations rather than causality; binary rewards on AIME restrict empirical return standard deviations to discrete values, reweighting does not flatten every checkpoint's trend, and the largest contributions are not consistently reduced. Noise normalization estimates weights from the same rollout group used to train the critic, which the authors note can introduce bias but works well in practice. The mini-batch analysis assumes fixed parameters and independent rollouts, while sequential Adam steps recompute gradients and update optimizer state, so the bound does not constrain the final parameter change. The seed study covers only FrontierCS with three runs each of EasyPPO and the actor-only-filtering baseline, while other tasks use one run per method; this evidence bundle is full text, but figures appear as references, so specific numbers should be checked against the original tables.
