acceptodds
Under review as a conference paper at ICLR 2027

FLORA: Dual-Latent Discrete World Models with One-step Continuous Flow Matching

Abstract

Discrete latent world models have demonstrated remarkable potential for sample-efficient visual reinforcement learning, enabling agents to learn policies from imagined rollouts in compact latent spaces. However, current methods face two key limitations. First, efficient one-step latent predictors typically use factorized sampling, predicting each token of the latent independently and ignoring inter-token dependencies; dependency-aware alternatives such as MaskGIT-based decoding improve coherence but require costly multi-step refinement during imagination. Second, reconstruction-based encoders preserve visual detail but may allocate capacity to task-irrelevant distractors, while reconstruction-free representations can lose fine-grained information needed for precise control. We propose **FLORA**, a discrete latent world model that addresses both limitations in a unified framework. FLORA introduces a *Denoising Latent Predictor (DP)* trained with continuous flow matching and shortcut forcing, enabling it to model inter-token dependencies in a single step with rollout cost comparable to factorized one-step predictors. It also introduces a *Dual-latent Representation (DR)* that pairs a stochastic reconstructive latent with an auxiliary feature-aligned latent, preserving visual detail while capturing task-relevant abstraction. Across Atari 100k, the DeepMind Control Suite, Meta-World, and Crafter, FLORA outperforms or matches strong model-based baselines. On Atari 100k, FLORA achieves a human-normalized mean score of 170% and a median score of 101%, making FLORA the first world model to surpass human-level median performance without decision-time planning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.