acceptodds
Under review as a conference paper at ICLR 2027

Streaming Flow Policy Optimization

Abstract

Proximal policy optimization (PPO) with a diagonal Gaussian actor is the standard method for massively parallel robot learning, yet its exploration is memoryless: independent action noise is injected at every control step. Physical behaviors, by contrast, unfold as temporally coherent trajectories, so exploration should be coherent as well. We introduce Streaming Flow Policy Optimization (SFPO), which builds temporal persistence into the actor while preserving PPO’s simplicity and efficiency. Flow and diffusion policies generate structured action trajectories, but their flow time indexes the noise level: policy gradients must score a K-step denoising chain or a surrogate, not the executed action. Streaming flow policies, designed for imitation learning, instead identify flow time with execution time, advancing the flow one deterministic Euler step per control step. SFPO replaces this step with an Euler–Maruyama step, so each control-step transition is an exact diagonal Gaussian over the executed action. Augmenting the state with the previous action preserves PPO’s exact per-step likelihood, importance ratio, entropy, and KL divergence at one network evaluation per control step. A single retention parameter interpolates between memoryless PPO and persistent streaming integration, giving a prior for correlated exploration without restricting the learned policy. On 13 of 14 continuous-control, manipulation, and locomotion tasks, with policies trained from scratch, SFPO matches or exceeds PPO and recent diffusion and flow-policy methods at PPO’s cost, with PPO’s hyperparameters and one fixed retention setting. These results position SFPO as a drop-in replacement for PPO with temporally coherent exploration.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.