acceptodds
Under review as a conference paper at ICLR 2027

Q-Guided Drifting Policy: Value-Guided Target Construction for One-Step Offline Reinforcement Learning

Abstract

Offline reinforcement learning (RL) often requires expressive stochastic policies to represent heterogeneous and multimodal behavior distributions. Recently proposed drifting models provide a promising generative modeling framework for offline RL policies, enabling multimodal action distributions with one-step stochastic generation. However, offline RL only provides behavior actions rather than the high-value target actions required to train drifting policies, creating a mismatch between drifting policy learning and offline supervision. To bridge this gap, we propose Q-Guided Drifting (QGD) policy, which uses a learned -function to construct training targets for drifting actors rather than directly optimizing the actor through critic gradients. QGD performs zeroth-order local refinement around multiple behavior samples, retaining high-value refined actions as positive targets and low-value behavior actions as additional negatives. The resulting policy preserves multiple high-value action modes while requiring only a single stochastic forward pass at inference. Experiments on OGBench and D4RL show that QGD achieves at least 95% of the best reported performance on 23 out of 37 tasks while retaining one-step stochastic inference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.