acceptodds
Under review as a conference paper at ICLR 2027

Disagreement in Imagination: A Cross-Timescale Consistency Signal for Robust Model-Based Control

Abstract

Planning inside a learned world model is sample-efficient but fragile: prediction errors compound with rollout depth, and execution-time disturbances cause plans to drift from reality. Every standard detector of model error such as filtering residuals, reconstruction losses requires an observation to compare against, and inside imagination no observation exists. We study a quantity that does not need one —the disagreement between two predictors of the same imagined step, one coarse and trained over abstract steps, one fine and trained over primitive steps. This cross-timescale disagreement is computable from predictions alone, needs no future observation, and is directional: it states how the fine predictor would revise the coarse one. It is consumed at three clocks: as a consistency penalty during planning, as a training regularizer that sharpens the abstract model, and, once the imagined step is realized, as the residual an online controller cancels. On 80 tasks across DMControl, Meta-World, and MyoSuite under combined observation and action noise, HiWM achieves the highest IQM among evaluated methods with a margin that grows monotonically with noise intensity. Ablations on a completed factorial confirm super-additive coupling between the feedback pathway and the controller. Freezing the fast predictor at its initial observation as the configuration closest to a standard posterior, loses half the benefit, and two prediction rates without upward flow perform worse than a single rate. The mechanism transfers zero-shot to a Franka Panda robot, where disagreement is 2.4 larger than in simulation and the controller's corrections scale proportionally.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.