Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

HDPO has LLMs list candidate solution outlines before selecting one to reason through, reportedly improving both reasoning and candidate-solution diversity

The work proposes Hint-Guided Diversified Policy Optimization (HDPO), which lets a model first list all potential candidate solution outlines as hints and then select the most reliable one for further reasoning; the method has two stages, Cold Start for Structured Reasoning and Hint-Guided Diversified Reinforcement Learning, to incentivize diverse and reliable solutions along a "propose-select-think" trajectory, and experimental results are reported to show that HDPO effectively boosts LLM reasoning and enhances the diversity of candidate solutions as well as the model's ability to identify reliable solutions.