Skip to main content
Back to timeline
arXivSource publication:

RRT turns rubric verdicts into item-response quality rewards, beating GRPO by 1.7 points on Qwen3.5-4B while cutting roughly half the judge requests

Synopsis

The work introduces Rubric Response Theory (RRT), which treats rubric criterion verdicts as item-response evidence about a shared latent quality, infers each rollout's quality as the GRPO reward via a two-parameter item response model, and uses a Response Parameter Network to predict criterion difficulty and discrimination from prompt and criterion text with online EM updates as the policy changes; with Qwen3.5-4B as the policy, RRT's macro criterion score across Medical, Science, Rubrics as Rewards Science, and RubricBench is 1.7 points above GRPO, gains 2.8 to 5.6 points on Hard and Very hard criteria in Medical and Science, and at half the criterion budget adaptive Fisher selection keeps the macro criterion score within 0.1 points of GRPO with full judging.

AI-generated editorial illustration: Rubric Rewards from Item Response Theory

Interpretation

RRT replaces summing assigned points for satisfied criteria with Bayesian inference of a latent quality, using the posterior mode as the GRPO reward, so rollouts with equal point totals but different verdict patterns can receive different rewards. Prior aggregation (e.g., Gunjal et al., Arora et al.) takes a weighted average of satisfied criteria using assigned points, where fixed points encode how much a criterion should count rather than how strongly its verdict distinguishes current rollouts; RRT instead explains the full verdict vector through difficulty and discrimination parameters. Theorem 2 shows that under the item response model the rubric likelihood score maximizes the local signal-to-noise ratio for quality among all scalar functions of the verdicts and reaches the Fisher information bound; Proposition 1 shows rollouts with equal point totals can still be separated. Empirically RRT has 1.2 to 2.2 times the within-prompt variance share of the points reward and 2% to 58% fewer tied pairs across 12 dataset and policy cells.

A Response Parameter Network predicts criterion difficulty and discrimination from prompt and criterion text, and online EM updates these estimates as the policy distribution changes, so parameters do not rely on fixed empirical pass rates. Earlier IRT work calibrates fixed examinees or models for assessment, data selection, and efficient evaluation; RRT applies this measurement to policy training, where each rollout's measured quality becomes its reward, and text prediction supports unseen rubrics. The online RPN improves macro Pearson correlation by 1.7 points on the next policy step (Medical +2.7, Science +0.7); predicted versus empirical difficulty has Spearman correlations of 38.0% to 47.9%; the hard E-step improves ROC-AUC by 0.6 to 1.3 points and reduces criterion loss by 0.011 to 0.017 over the tested soft variants.

The same Fisher information serves both reward inference and criterion selection, reducing judge requests under a criterion budget. Adaptive testing has long selected questions by information, but this had not been tied to rubric criterion selection during policy training; RRT ranks unjudged criteria at the currently inferred qualities and updates those qualities after each judged criterion. On rollouts from trained policies, adaptive Fisher selection reaches 95.0% mean Pearson correlation while leaving 21.0% of criteria unjudged, 11.0 points more than random selection; at criterion budget 0.50 RRT's macro criterion score stays within 0.1 points of GRPO with full judging, judge requests fall 49.0% on Medical and Science, median judging time falls 49.1% to 49.6%, and total step time falls 23.5% to 28.7%.

In the primary Qwen3.5-4B comparison, RRT with the online RPN is highest or tied in every column of both macro metrics and gains most on difficult criteria. Relative to Vanilla GRPO, POW3R, and DIVA, which also adapt aggregation to current rollouts, RRT replaces the weighted sum and additionally provides criterion selection. RRT with the online RPN exceeds Vanilla GRPO by 1.7 points and POW3R by 0.8 points; it beats Vanilla GRPO in seven of eight difficulty bands on Medical and Science, including every Medium, Hard, and Very hard band, by 2.8 to 5.6 points; it gains 0.1 to 0.7 macro criterion points on HealthBench and ResearchQA; and it has lower median response length in all 12 dataset and policy combinations.

Perspective

The result targets rubrics whose criteria are monotone indicators of one shared quality, used when training language model policies with GRPO-style algorithms for reward computation and criterion selection; the authors state that RRT complements rather than replaces rubric generation and note that multiple quality targets could accommodate rubrics with explicit tradeoffs. For practitioners this means that when judging cost is the main bottleneck and rubric criteria share one quality target, text-predicted difficulty and discrimination can replace fixed points, and the most informative criteria can be selected under a criterion budget; the trained policy adds no parameters or inference components at deployment.

Whether criteria truly measure one shared quality remains open: the authors ask in the community discussion about rubrics whose criteria do not all measure one shared quality, and diagnostics show fitted quality explains 72.8% to 85.3% of pairwise mutual information while the most similar criterion-text band has 3.4 to 11.9 points more redundancy than the bootstrap null. The primary comparison uses one training run per condition, and the reported intervals quantify variation across prompt groups rather than training randomness. On Qwen3.5-2B and Llama-3.1-8B-Instruct the macro criterion scores are close to Vanilla GRPO (0.2 points below and 0.1 points above), so robustness of the gain across policies and runs remains to be confirmed. The noise-structure analysis also shows criterion score leading by 0.8 to 1.2 points when noise is tied to response length or symmetric across all criteria, indicating RRT's advantage depends on judge errors concentrating on particular criteria.

Sources