acceptodds
Under review as a conference paper at ICLR 2027

Hint-Guided Diversified Policy Optimization for LLM Reasoning

Abstract

Recent developments in Large Language Models (LLMs) have showcased impressive reasoning capabilities and a promising enhancement strategy is Reinforcement Learning with Verifiable Rewards (RLVR). However, existing reward mechanisms are constrained to the outcome-level correctness and lack explicit signals to guide the model to consider diverse solutions. In contrast, human problem-solving usually involves considering several potential solutions and choosing the most reliable one, a cognitive process that current RLVR frameworks do not explicitly incentivize. Inspired by this cognitive process, we propose Hint-Guided Diversified Policy Optimization (HDPO), allowing the model to first list all potential candidate solution outlines as hints and then select the most reliable one for further reasoning. HDPO comprises two stages of **Cold Start for Structured Reasoning** and **Hint-Guided Diversified Reinforcement Learning** to incentivize the model to generate diverse and reliable solutions following the “propose-select-think” trajectory. Experimental results show that HDPO effectively boosts LLM reasoning and enhances the diversity of candidate solutions as well as the LLM's ability to identify reliable solutions across mathematical and knowledge-intensive tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.