acceptodds
Under review as a conference paper at ICLR 2027

Reinforced Planning with Latent World Models

Abstract

Humans solve complex problems by constructing plans and mentally simulating their outcomes with an internal model of the world. Machine learning has made substantial progress in learning world models that predict the consequences of action sequences, yet the procedures used to plan with these models remain largely hand-designed. Most planners rely on fixed search or optimization rules; approaches that learn aspects of search typically imitate a predefined optimizer or use planning to train an amortized policy, rather than learning how to iteratively improve a complete multi-step plan. We introduce Reinforced Planning, a method for learning the plan-update rule itself through a pretrained world model. Our implementation, RP1, learns a critic for imagined outcomes by temporal-difference learning, and a neural operator that iteratively revises multi-step action sequences, trained by differentiating the critic's value through the world model. RP1 requires no optimizer demonstrations and is trained without environment interaction, from imagined rollouts of a pretrained differentiable latent world model; environment episodes are used only for checkpoint selection. Across visual navigation, arm reaching, and robotic manipulation on two world-model backbones, RP1 matches or exceeds hand-designed and learned planners, achieving near-perfect success in several settings while using fewer world-model rollouts than the strongest alternative (CEM) and planning up to faster under concurrent planners inference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.