acceptodds
Under review as a conference paper at ICLR 2027

ShiftWM-Bench: Planning with Frozen Latent World Models under Hidden Within-Episode Dynamics Shifts

Abstract

Pretrained latent world models are typically evaluated under fixed within-episode dynamics, although changes during execution can invalidate the action-outcome relationships used for planning. We introduce ShiftWM-Bench, a benchmark of four task-specific, state-triggered dynamics changes in manipulation and navigation, with event timing and response parameters hidden from the planner. We study a planning stack with two complementary components and no test-time parameter updates. Proposal Shift combines latent retrieval, cross-entropy method (CEM) initialization, and reliability-weighted action-prior scoring to strengthen candidate search. Latent Belief Recovery infers a distribution over predefined response hypotheses from recent prediction-observation residuals, allocates candidate plans according to this distribution, and ranks them under the most probable hypothesis. The belief model is trained on nominal trajectories and synthetic response perturbations, without real shifted training trajectories. Across 300 scenarios per task and three planner seeds per scenario, Proposal Shift raises four-task mean success from 47.7% with native LeWM to 64.8%, and Recovery further raises it to 72.7%. Controlled post-trigger success gains are established on Cube under particular handoff protocols and control budgets; command-delay gains remain trajectory-dependent, and the current contact-friction approximation can degrade planning. These evaluations measure candidate-search improvements and conditional post-shift benefit separately, characterizing the operating boundary of the complete response-conditioned planner.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.