acceptodds
Under review as a conference paper at ICLR 2027

RPO: Routing Policy Optimization for Online Diffusion Reinforcement Learning

Abstract

Reinforcement learning can align visual generators with reward signals beyond their pretraining objectives, but existing score-function methods typically introduce stochasticity by perturbing the reverse denoising trajectory, leading to costly rollout-based optimization. We instead ask whether stochastic exploration can be introduced inside the generator while preserving its original sampler. Routing Policy Optimization (RPO) answers this question by formulating attention routing as a discrete policy. At a selected denoising level, each query block stochastically attends to a subset of candidate key blocks, with Bernoulli probabilities derived from the generator’s native query–key affinities. The resulting routing masks define discrete actions that can be optimized with a score-function objective. RPO further uses a controlled two-rollout estimator: two rollouts share the same prompt, initial latent, routed denoising level, and routed layers, differing only in their sampled attention routes. Their reward difference provides an advantage signal for the routing actions. Because the routing policy reuses the generator’s native attention parameters, the learned updates remain effective after routing is removed. At inference, RPO restores dense attention and retains the original deterministic sampler. Experiments on text-to-image and text-to-video generation show that RPO improves generation quality while achieving nearly 4× higher training efficiency than Flow-GRPO-style methods. More details and results are available at https://anonymous.4open.science/r/RoutePO-submission.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.