Extracting Behavior from Frozen Latent World Models: Experience-Anchored Proposals with Offline Verifier Alignment
Abstract
Planning in frozen latent world models is usually performed with the cross-entropy method (CEM). CEM is memoryless: every replan starts a new Gaussian search that uses none of the experience the model was trained on. Planning is therefore slow, 2.1s per replan in our setup, and uneven, from 67% to 91% success across four suites. We replace this search with a proposer that learns from the model's experience. A generative proposer is behavior-cloned from the corpus. At plan time it generates candidates and the model selects among them by imagined latent distance to the goal. We then improve the proposer by offline reinforcement learning with this imagined distance as the reward, and analytic relative-position constraints can be added at plan time without retraining. We recommend a rectified-flow proposer, the LeWM Action Proposer (LeAP), improved by World-model Advantage-weighted Verifier Regression (WAVR), which needs no trust-region term. With 256 candidates, a replan takes 24ms and of CEM's model evaluations, about 90 times faster per planning call. On Cube, TwoRoom and PushT, whose corpora contain directed or expert behavior, the behavior-cloned proposer alone reaches 99.9%, 98.0% and 96.2%, against 67.0%, 88.8% and 91.4% for CEM. On Reacher, whose corpus is random behavior, WAVR raises success from 76.1% to 97.3%, against 84.3% for CEM. With a keep-out region added on Cube at plan time, the planner reaches the goal without entering it in 91.6% of episodes, against 55.9% for CEM.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.