Public articles linked to the same research event.
arXiv The work proposes Hint-Guided Diversified Policy Optimization (HDPO), which lets a model first list all potential candidate solution outlines as hints and then select the most reliable one for further reasoning; the method has two stages, Cold Start for Structured Reasoning and Hint-Guided Diversified Reinforcement Learning, to incentivize diverse and reliable solutions along a "propose-select-think" trajectory, and experimental results are reported to show that HDPO effectively boosts LLM reasoning and enhances the diversity of candidate solutions as well as the model's ability to identify reliable solutions.
The work proposes Hint-Guided Diversified Policy Optimization (HDPO), which lets a model first list all potential candidate solution outlines as hints and then select the most reliable one for further reasoning; the method has two stages, Cold Start for Structured Reasoning and Hint-Guided Diversified Reinforcement Learning, to incentivize diverse and reliable solutions along a "propose-select-think" trajectory, and experimental results are reported to show that HDPO effectively boosts LLM reasoning and enhances the diversity of candidate solutions as well as the model's ability to identify reliable solutions.
The work proposes Hint-Guided Diversified Policy Optimization (HDPO), which lets a model first list all potential candidate solution outlines as hints and then select the most reliable one for further reasoning; the method has two stages, Cold Start for Structured Reasoning and Hint-Guided Diversified Reinforcement Learning, to incentivize diverse and reliable solutions along a "propose-select-think" trajectory, and experimental results are reported to show that HDPO effectively boosts LLM reasoning and enhances the diversity of candidate solutions as well as the model's ability to identify reliable solutions.
The work proposes Hint-Guided Diversified Policy Optimization (HDPO), which lets a model first list all potential candidate solution outlines as hints and then select the most reliable one for further reasoning; the method has two stages, Cold Start for Structured Reasoning and Hint-Guided Diversified Reinforcement Learning, to incentivize diverse and reliable solutions along a "propose-select-think" trajectory, and experimental results are reported to show that HDPO effectively boosts LLM reasoning and enhances the diversity of candidate solutions as well as the model's ability to identify reliable solutions.