acceptodds
Under review as a conference paper at ICLR 2027

Learned DRIFT on the Logit Manifold for LLM Reasoning Exploration

Abstract

Reinforcement learning from verifiable rewards (RLVR) can only reinforce what it samples, which makes the rollout group the bottleneck: the policy can progress no further than the spread of rewards in the group allows, and only along the trajectories it has already sampled. Existing exploration methods push on both axes blindly, temperature scaling, entropy regularization, and looser clipping widening the sampling distribution in every direction at once while group filtering simply samples harder. We introduce **DRIFT** (**D**irected **R**ollout with **I**nferred **F**eedback at **T**riggers), which applies a Langevin step to steer generation itself, to produce groups with diverse rewards and trajectories. The optimal drift is the gradient of expected reward in logit space, an expectation over completions that no rollout can evaluate before it terminates. With a first-order expansion, we show that a step along any partially aligned direction earns the optimal drift's gain scaled by the cosine between the two. DRIFT therefore estimates such a direction online, from prior steering feedback, using the entropy remaining at branch points as its signal. Training on Qwen3-4B-Base and Qwen3-8B-Base, DRIFT improves reasoning performance under both GRPO and DAPO, raising the benchmark average by 1.5 to 3.1 points in every pairing and AIME accuracy by up to 4.6 points.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.