Public articles linked to the same research event.
arXiv The work proposes a two-stage training framework in which a student without rubric access first matches a rubric-aware teacher's next-token distributions at student-generated prefixes (RP-OPD), then optimizes the rubric reward directly with RL; across HealthBench, ResearchQA, and RubricHub Science, comparing post-training methods and varying the amount of SFT or RP-OPD before RL, the two-stage framework achieves the highest scores among the methods evaluated, and RP-OPD + RL shows limited signs of reward hacking on RubricHub Science whereas the SFT + RL baseline increasingly receives high rewards for claims of rubric compliance without providing the required content.
The work proposes a two-stage training framework in which a student without rubric access first matches a rubric-aware teacher's next-token distributions at student-generated prefixes (RP-OPD), then optimizes the rubric reward directly with RL; across HealthBench, ResearchQA, and RubricHub Science, comparing post-training methods and varying the amount of SFT or RP-OPD before RL, the two-stage framework achieves the highest scores among the methods evaluated, and RP-OPD + RL shows limited signs of reward hacking on RubricHub Science whereas the SFT + RL baseline increasingly receives high rewards for claims of rubric compliance without providing the required content.
The work proposes a two-stage training framework in which a student without rubric access first matches a rubric-aware teacher's next-token distributions at student-generated prefixes (RP-OPD), then optimizes the rubric reward directly with RL; across HealthBench, ResearchQA, and RubricHub Science, comparing post-training methods and varying the amount of SFT or RP-OPD before RL, the two-stage framework achieves the highest scores among the methods evaluated, and RP-OPD + RL shows limited signs of reward hacking on RubricHub Science whereas the SFT + RL baseline increasingly receives high rewards for claims of rubric compliance without providing the required content.
The work proposes a two-stage training framework in which a student without rubric access first matches a rubric-aware teacher's next-token distributions at student-generated prefixes (RP-OPD), then optimizes the rubric reward directly with RL; across HealthBench, ResearchQA, and RubricHub Science, comparing post-training methods and varying the amount of SFT or RP-OPD before RL, the two-stage framework achieves the highest scores among the methods evaluated, and RP-OPD + RL shows limited signs of reward hacking on RubricHub Science whereas the SFT + RL baseline increasingly receives high rewards for claims of rubric compliance without providing the required content.