acceptodds
Under review as a conference paper at ICLR 2027

PoleWM: Visual World Models with Pole-Based Action Responses

Abstract

Visual planning requires predicting how candidate actions unfold over time. Direct world models avoid recursive state rollout, but direct prediction alone does not specify how action effects should share temporal structure. We introduce PoleWM, which represents a latent trajectory as an initial-condition response plus delayed, action-conditioned responses. Nonlinear networks determine response coefficients from visual history and causal action prefixes; a shared pole basis determines their evolution with elapsed time. This factorization preserves nonlinear action interactions while providing exact initial-state anchoring, non-anticipation, and bounded responses to finite plans. We also characterize the additional basis direction needed to represent re-anchored tails, separating these structural properties from restart consistency. Across four image-goal control environments, PoleWM achieves 90.83% mean success over two training seeds, exceeding the strongest baseline mean by 6.08 percentage points and leading on three tasks. In a first-epoch Cube study with one training seed and no restart penalty, neural and pole response heads improve over a generic direct head by 20.67 and 23.34 points, respectively. Separately, optimized single-frame planning takes 74-87 ms per complete CEM solve on an A100. These results support explicit action-response factorization as an effective inductive bias for direct visual world models, with poles providing an analytically structured realization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.