Skip to main content
Back to timeline
arXivSource publication:

Range-GRPO compares conformally calibrated reward intervals pairwise, achieving the highest in-distribution and out-of-distribution average performance among evaluated semi-supervised methods with fewer training resources

Synopsis

The work proposes Range-GRPO, a semi-supervised post-training framework that represents LLM-as-a-Judge pseudo-rewards as conformally calibrated reward ranges and compares reward ranges pairwise within GRPO rather than reducing them to point rewards, letting interval uncertainty affect both the magnitude and direction of learning signals; theoretical analysis shows the objective recovers the GRPO advantage when all reward ranges collapse to points, and empirically Range-GRPO achieves the highest in-distribution and out-of-distribution average performance among the evaluated semi-supervised methods while requiring fewer training resources.

Source-provided article image: Range-GRPO: Policy Optimization via Pairwise Relations among Reward Intervals
Figure 1 ·

Figure 1: Overview of Range-GRPO and its learning signals. (a) At initialization, Range-GRPO fits a state scorer on labeled source rollouts and computes nonconformity scores on calibration rollouts. Each iteration generates policy and target DRE rollouts with G G and G ′ G^{\prime} responses, respectively, and updates density ratios using the target states and reused source states. The density ratios, fixed scorer, and calibration scores are used to construct reward ranges, from which pairwise interval comparisons construct the advantages for policy updates. (b) An illustrative comparison uses identical reward intervals for both methods. Although the interval for o 2 o_{2} lies above that for o 3 o_{3} with the same width, midpoint reduction assigns it a negative advantage under Dr.GRPO. Range-GRPO instead downweights pairwise contributions involving wider, more uncertain intervals, turning this advantage positive and aligning the direction of the learning signal with the oracle preference.

arXiv

Interpretation

It proposes Range-GRPO, which represents pseudo-rewards as conformally calibrated reward ranges and compares reward ranges pairwise within GRPO instead of first reducing them to point rewards. Relative to LLM-as-a-Judge practice that uses a single point score as the pseudo-reward, this objective lets interval uncertainty affect both the magnitude and direction of the learning signal. The abstract provides a method-level description and a theoretical analysis showing the objective recovers the GRPO advantage when all reward ranges collapse to points; concrete experimental settings and numbers are not given in the abstract.

It builds a semi-supervised post-training framework that combines limited labeled data with unlabeled prompts. For domains without reference answers or executable verifiers, it combines scalable pseudo-rewards with a small amount of reliable supervision rather than relying on only one of them. The abstract states the framework combines limited labeled data with unlabeled prompts and reports the best performance among the evaluated semi-supervised methods; data scale and labeling ratio are not provided.

Among the evaluated semi-supervised methods, Range-GRPO achieves the highest in-distribution and out-of-distribution average performance while requiring fewer training resources. Incorporating interval uncertainty into GRPO's relative comparisons is reported to lead on both in-distribution and out-of-distribution evaluation averages, together with lower training resource requirements. The abstract reports this comparison result but does not list baselines, evaluation tasks, metric values, or how resources were measured.

Perspective

The work targets post-training settings that lack reference answers or executable verifiers and therefore rely on LLM-as-a-Judge pseudo-rewards, and it applies to semi-supervised settings where a small amount of labeled data and many unlabeled prompts are available. Its objective design brings reward-range uncertainty into GRPO's relative comparisons, making it directly relevant to researchers and practitioners who want to use unlabeled prompts without increasing labeling cost; the theoretical analysis states its applicability condition as recovering the GRPO advantage when all reward ranges collapse to points.

The abstract does not list the compared semi-supervised methods, evaluation tasks, or metric values, nor how training resources were measured, so the size of the performance lead and the degree of resource savings still need to be confirmed in the full text. The specific construction of the conformal calibration, how interval width affects optimization stability, and the assumptions behind the theoretical analysis are open questions a reader should keep in mind when reading the full paper.

Sources