acceptodds
Under review as a conference paper at ICLR 2027

Scaling World-Model Reinforcement Learning Through Diffusion Policy Optimization

Abstract

Model-based reinforcement learning (RL) can be effectively supported at scale through the use of world models. However, scaling such approaches remains fundamentally limited because policy improvement often relies on value functions induced by a separate non-search policy, creating training inconsistency and ultimately leading to suboptimal learning. To address this limitation, we propose Model-Based Diffusion Policy Optimization (MBDPO) in world models, a framework that unifies search and policy optimization through diffusion policy representations. Instead of constructing an explicit planner over a learned world model, we reformulate policy optimization as a diffusion process over searched trajectories in latent world models. In this view, we extract an implicit energy function from the collected transitions that anchors the policy, enabling MBDPO to refine the score field for policy optimization while mitigating misalignment. We evaluate MBDPO across several complementary settings, including offline pretraining, online learning, offline-to-online fine-tuning, and massively multi-task online learning. MBDPO achieves strong multi-task performance and exhibits consistent scaling behavior with increasing model size.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.