Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

RLCPR characterizes posterior concentration and beats a state-of-the-art verifier-free RL baseline by up to 4.0% on six of seven benchmarks

This work identifies a length-dependent failure mode of probability-based rewards in verifier-free reinforcement learning, termed the Posterior Concentration Phenomenon (PCP), in which the probability of a reference answer conditioned on a reasoning trace collapses to a low-variance interval as the trace grows long and rewards become nearly indistinguishable, and proposes RLCPR, a framework with uncertainty-aware data sampling and concentration-aware regularization that improves optimization stability and token efficiency, outperforming the state-of-the-art verifier-free RL baseline by up to 4.0% on six of seven benchmarks.