Public articles linked to the same research event.
arXiv The authors introduce CATCH, a controllable testbed for reward hacking in coding RL that deliberately exposes environmental loopholes and provides execution-based gold labels by comparing success under a vulnerable evaluator with task correctness under an independent audit, while controlling the model's initial hacking tendency through supervised fine-tuning data mixtures and the difficulty of earning rewards through reward designing; experiments show CATCH produces diverse RL training trajectories with clear reward hacking, that initial models and reward difficulties jointly shape its emergence, and that a chain-of-thought monitor initially suppresses hacking but this protection erodes as the policy model learns to mislead the monitor with code comments.
The authors introduce CATCH, a controllable testbed for reward hacking in coding RL that deliberately exposes environmental loopholes and provides execution-based gold labels by comparing success under a vulnerable evaluator with task correctness under an independent audit, while controlling the model's initial hacking tendency through supervised fine-tuning data mixtures and the difficulty of earning rewards through reward designing; experiments show CATCH produces diverse RL training trajectories with clear reward hacking, that initial models and reward difficulties jointly shape its emergence, and that a chain-of-thought monitor initially suppresses hacking but this protection erodes as the policy model learns to mislead the monitor with code comments.
The authors introduce CATCH, a controllable testbed for reward hacking in coding RL that deliberately exposes environmental loopholes and provides execution-based gold labels by comparing success under a vulnerable evaluator with task correctness under an independent audit, while controlling the model's initial hacking tendency through supervised fine-tuning data mixtures and the difficulty of earning rewards through reward designing; experiments show CATCH produces diverse RL training trajectories with clear reward hacking, that initial models and reward difficulties jointly shape its emergence, and that a chain-of-thought monitor initially suppresses hacking but this protection erodes as the policy model learns to mislead the monitor with code comments.
The authors introduce CATCH, a controllable testbed for reward hacking in coding RL that deliberately exposes environmental loopholes and provides execution-based gold labels by comparing success under a vulnerable evaluator with task correctness under an independent audit, while controlling the model's initial hacking tendency through supervised fine-tuning data mixtures and the difficulty of earning rewards through reward designing; experiments show CATCH produces diverse RL training trajectories with clear reward hacking, that initial models and reward difficulties jointly shape its emergence, and that a chain-of-thought monitor initially suppresses hacking but this protection erodes as the policy model learns to mislead the monitor with code comments.