acceptodds
Under review as a conference paper at ICLR 2027

BlendPO: Controlling Guidance Strength in Reinforcement Learning without Changing Guidance Content

Abstract

Reinforcement learning with verifiable rewards provides no learning signal when sampled rollouts fail to produce a correct solution. A common approach is to condition rollouts on guidance, such as partial solutions or hints, during training. Stronger guidance makes successful rollouts more likely but can transfer less to the unguided setting used at test time, making guidance strength a key design choice. Solution prefixes allow fine-grained control by adjusting the prefix length. In contrast, for guidance without a natural ordering, such as abstract hints, the guidance strength is typically controlled only through coarse choices, such as selecting a hint level or deciding whether to provide a hint at all. We propose BlendPO (Blended-context Policy Optimization), which enables fine-grained control of guidance strength without changing guidance content by interpolating policy logits obtained with and without guidance. Rather than training with a single guidance strength, BlendPO samples responses at multiple interpolation strengths for each problem and updates the policy using these responses under the original unguided context. We first evaluate BlendPO on mathematical reasoning using abstract hints generated from reference solutions. BlendPO outperforms methods based on solution prefixes as well as those using the same abstract hints. We then extend BlendPO to visual reasoning with visual states, textual descriptions of the visual input that can be obtained from dataset metadata or the task environment without reference solutions. Across four visual reasoning task families, BlendPO consistently improves over GRPO and methods using the same visual states, demonstrating its potential to broaden the practical applicability of guided RL when reference solutions are unavailable.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.