Taming Generative Control Policies for Reinforcement Learning
Abstract
Generative control policies have emerged as an effective policy parameterization for robot learning. They leverage generative models such as diffusion and flow matching to iteratively transform noise into actions, enabling these policies to effectively model complex action distributions. However, prior work has reported that optimizing these policies with reinforcement learning (RL) can lead to training instability. This instability has often been associated with the iterative nature of the generation process. Many existing methods therefore avoid directly backpropagating value gradients through the entire iterative generation process, leading suboptimal performance. Contrary to prior belief, We show that iterative generation can support stable value-gradient RL and identify sampler design as an important factor in optimization. Motivated by this insight, we propose VINE, a new sampling strategy that reconstructs the generation path at each iteration, enabling stable end-to-end value-gradient propagation. VINE requires only a few lines of changes to the vanilla flow matching sampler. VINE retains iterative generation and end-to-end value-gradient optimization through the entire sampling chain. Experiments show that VINE outperforms the evaluated full-BPTT baselines in aggregate on OGBench tasks. VINE enables stable off-policy RL adaptation of the pretrained VLA across LIBERO suites. On real-world manipulation tasks, VINE rapidly improves success rate toward 100% from scratch within 20 minutes of online rollouts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.