acceptodds
Under review as a conference paper at ICLR 2027

Taming Generative Control Policies for Reinforcement Learning

Abstract

Generative control policies have emerged as an effective policy parameterization for robot learning. They leverage generative models such as diffusion and flow matching to iteratively transform noise into actions, enabling these policies to effectively model complex action distributions. However, prior work has reported that optimizing these policies with reinforcement learning (RL) can lead to training instability. This instability has often been associated with the iterative nature of the generation process. Many existing methods therefore avoid directly backpropagating value gradients through the entire iterative generation process, leading suboptimal performance. Contrary to prior belief, We show that iterative generation can support stable value-gradient RL and identify sampler design as an important factor in optimization. Motivated by this insight, we propose VINE, a new sampling strategy that reconstructs the generation path at each iteration, enabling stable end-to-end value-gradient propagation. VINE requires only a few lines of changes to the vanilla flow matching sampler. VINE retains iterative generation and end-to-end value-gradient optimization through the entire sampling chain. Experiments show that VINE outperforms the evaluated full-BPTT baselines in aggregate on OGBench tasks. VINE enables stable off-policy RL adaptation of the pretrained VLA across LIBERO suites. On real-world manipulation tasks, VINE rapidly improves success rate toward 100% from scratch within 20 minutes of online rollouts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.