Adapting World Models through Response-Preserving Representations
Abstract
A world model can predict the next observation almost perfectly and still be the wrong model to plan with, because deployment physics rarely match the simulator it was trained in. A controller must read the change from a short prefix of frames and actions, then apply the same correction to every future it imagines. We show that ordinary forward training quietly undermines the first step: a history-dependent predictor explains part of a shifted successor from the history itself, absorbing the change and leaving too little residual for any calibration procedure to read. Forward and controlled-pair objectives cannot tell this allocation from an informative one, since models that divide the change differently make identical predictions. Response-preserving representations (RPR) fix the allocation rather than the symptom: the encoder keeps a grid of spatial motion tokens, controlled source pairs anchor an explicit response channel, and encoder, predictor, and response are trained through the very weighted ridge solve that runs at deployment. With the networks frozen, one solve on the prefix returns the change coefficient and a calibrated covariance; the coefficient is reused at every imagined step and refreshed as observations arrive. A linear-Gaussian analysis ties retained response to prefix estimation risk. On three simulated visual-control tasks, RPR recovers the change from eight transitions in under nine milliseconds, matches twenty-step gradient adaptation on the all-profile mean, and surpasses it wherever the change lies in the source family, while true-coefficient controls and absorption diagnostics attribute the gain to a better estimate and a better forward model alike.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.