Public articles linked to the same research event.
arXiv This work identifies a length-dependent failure mode of probability-based rewards in verifier-free reinforcement learning, termed the Posterior Concentration Phenomenon (PCP), in which the probability of a reference answer conditioned on a reasoning trace collapses to a low-variance interval as the trace grows long and rewards become nearly indistinguishable, and proposes RLCPR, a framework with uncertainty-aware data sampling and concentration-aware regularization that improves optimization stability and token efficiency, outperforming the state-of-the-art verifier-free RL baseline by up to 4.0% on six of seven benchmarks.
This work identifies a length-dependent failure mode of probability-based rewards in verifier-free reinforcement learning, termed the Posterior Concentration Phenomenon (PCP), in which the probability of a reference answer conditioned on a reasoning trace collapses to a low-variance interval as the trace grows long and rewards become nearly indistinguishable, and proposes RLCPR, a framework with uncertainty-aware data sampling and concentration-aware regularization that improves optimization stability and token efficiency, outperforming the state-of-the-art verifier-free RL baseline by up to 4.0% on six of seven benchmarks.
This work identifies a length-dependent failure mode of probability-based rewards in verifier-free reinforcement learning, termed the Posterior Concentration Phenomenon (PCP), in which the probability of a reference answer conditioned on a reasoning trace collapses to a low-variance interval as the trace grows long and rewards become nearly indistinguishable, and proposes RLCPR, a framework with uncertainty-aware data sampling and concentration-aware regularization that improves optimization stability and token efficiency, outperforming the state-of-the-art verifier-free RL baseline by up to 4.0% on six of seven benchmarks.
This work identifies a length-dependent failure mode of probability-based rewards in verifier-free reinforcement learning, termed the Posterior Concentration Phenomenon (PCP), in which the probability of a reference answer conditioned on a reasoning trace collapses to a low-variance interval as the trace grows long and rewards become nearly indistinguishable, and proposes RLCPR, a framework with uncertainty-aware data sampling and concentration-aware regularization that improves optimization stability and token efficiency, outperforming the state-of-the-art verifier-free RL baseline by up to 4.0% on six of seven benchmarks.