acceptodds
Under review as a conference paper at ICLR 2027

One Flow Transformer for Imagination and Control

Abstract

Diffusion and flow models are effective world models for visual reinforcement learning, letting agents learn behaviors in imagination. Yet existing diffusion-based online agents use them as black-box simulators: a separate policy network learns from generated rollouts, leaving the backbone’s representations unused for control. We introduce DRIFT, an efficient agent whose policy learns entirely in imagination on the same Flow Transformer backbone as its world model, trained online from scratch. DRIFT’s imagination is fast: it runs in a compact continuous latent space, using 33× fewer FLOPs per imagined transition than prior diffusion-based agents. On online RL benchmarks, DRIFT outperforms diffusion and flow baselines on Craftium and Crafter, and is competitive on Atari and with discrete latent-dynamics agents. Three modeling choices make this possible. First, we find that denoising alone yields a less effective control representation. We therefore train the backbone to predict the next latent from the policy’s compact state, a reconstruction-free objective that improves control. Second, we observe that reward predictions are less accurate on generated latents than on real ones, which can hinder policy learning. A simple and effective mitigation is to also train the world model on its own rollouts. Third, shortcut flow matching reduces imagination to one denoising step per transition, matching multi-step sampling in rollout accuracy. These results show that one generative backbone can both imagine and control in online RL.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.