Learning Intrinsic Dynamics in Latent World Models
Abstract
Latent world models can suffer substantial planning degradation when visual appearances vary, even when the underlying physical trajectories remain unchanged. Planning requires preserving intrinsic dynamics, the transition structure that links actions to their physical consequences. Yet latent prediction alone need not penalize the loss of distinctions between those consequences. We study action predictability as an additional constraint: consecutive latent states should retain information about the executed action. For distinct actions from the same state under a fixed appearance, we prove a positive lower bound on inverse prediction error when distinguishable outcomes are encoded identically. Forward prediction can still match their shared latent target exactly on these transitions. We implement this constraint with Concat IDM, an auxiliary inverse dynamics head that predicts actions from concatenated consecutive latent states and whose loss updates the shared encoder using existing offline action labels. The auxiliary head is discarded at test time. Across three independent training runs on Push-T with 24 appearances, Concat IDM increases planning success from to . The gain over the baseline is percentage points larger with 24 appearances than with a single appearance. The improvement persists across independently sampled assignments of appearances to training episodes under a common evaluation. On TwoRoom, Concat IDM increases success from to . These results support action predictability as an inductive bias for preserving intrinsic dynamics under visual diversity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.