QPILOTS: Efficient Test-Time Q-Steering for Flow Policies
Abstract
Flow-matching and diffusion policies are expressive action generators, but optimizing them with reinforcement learning (RL) is difficult: backpropagating the critic's action gradient through a multi-step denoising process is unstable, so existing methods distill the policy into a one-step actor or repeatedly fine-tune it as the critic improves. We show that inference-time steering alone, without ever training the policy against the critic, can outperform these fine-tuning approaches. Our method, QPILOTS, steers each denoising step of a flow policy: it maps the noisy intermediate action to an estimate of the clean action, where the critic is reliable, and adds the critic gradient computed there, rescaled to the magnitude of the base velocity. On a standard offline-to-online RL benchmark, QPILOTS achieves the best aggregate performance across 50 tasks, above baselines that fine-tune the policy against the critic. We also steer a frozen, pretrained VLA model, outperforming prior inference-time approaches in aggregate over six simulated manipulation tasks and on one real-robot task. We ablate key design choices against practicality and performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.