Q-Guided Drifting Policy: Value-Guided Target Construction for One-Step Offline Reinforcement Learning
Abstract
Offline reinforcement learning (RL) often requires expressive stochastic policies to represent heterogeneous and multimodal behavior distributions. Recently proposed drifting models provide a promising generative modeling framework for offline RL policies, enabling multimodal action distributions with one-step stochastic generation. However, offline RL only provides behavior actions rather than the high-value target actions required to train drifting policies, creating a mismatch between drifting policy learning and offline supervision. To bridge this gap, we propose Q-Guided Drifting (QGD) policy, which uses a learned -function to construct training targets for drifting actors rather than directly optimizing the actor through critic gradients. QGD performs zeroth-order local refinement around multiple behavior samples, retaining high-value refined actions as positive targets and low-value behavior actions as additional negatives. The resulting policy preserves multiple high-value action modes while requiring only a single stochastic forward pass at inference. Experiments on OGBench and D4RL show that QGD achieves at least 95% of the best reported performance on 23 out of 37 tasks while retaining one-step stochastic inference.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.