RetraQ: Manifold-Constrained Critic Guidance for Flow Policies in RL
Abstract
A learned critic can improve a pretrained flow-matching policy by steering sampled actions toward higher predicted value. However, even a critic that fits the behavior data well can provide misleading action gradients. Gradient components pointing away from the data-supported region are weakly constrained by behavior data and can lead guidance to exploit critic extrapolation errors. We introduce RetraQ (Retraction-based Q-guidance), which uses geometry encoded by the pretrained flow to constrain both the direction and the destination of critic guidance. A single velocity evaluation of the frozen flow defines a denoising map. RetraQ filters the critic gradient with the map's Jacobian and applies the map itself to correct the updated action. Our analysis characterizes when these operations approximate tangent filtering and retraction onto the behavior manifold. The same update guides actions at test time and provides targets for policy distillation and online reinforcement learning. In controlled studies, RetraQ improves true return while keeping refined actions close to the data-supported region, and its distillation targets remain more stable than those of raw critic ascent. In offline-to-online training on LIBERO and OGBench, RetraQ achieves higher mean final success than the evaluated critic-guidance baselines. On a real-world plug-insertion task, test-time RetraQ improves the success rate of a behavior-cloned policy from % to %. Online training further increases success to % in only 30 minutes and % within one hour. These results show that a pretrained flow policy provides not only an expressive action distribution, but also local geometry for reliable policy improvement.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.