Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Distilling on-policy with rubrics before RL tops the evaluated methods on health and science tasks

The work proposes a two-stage training framework in which a student without rubric access first matches a rubric-aware teacher's next-token distributions at student-generated prefixes (RP-OPD), then optimizes the rubric reward directly with RL; across HealthBench, ResearchQA, and RubricHub Science, comparing post-training methods and varying the amount of SFT or RP-OPD before RL, the two-stage framework achieves the highest scores among the methods evaluated, and RP-OPD + RL shows limited signs of reward hacking on RubricHub Science whereas the SFT + RL baseline increasingly receives high rewards for claims of rubric compliance without providing the required content.