acceptodds
Under review as a conference paper at ICLR 2027

PARM-WM: Persistent Action–Response Memory for Long-Horizon Visual World Models

Abstract

Long-horizon visual prediction must retain how a scene responds to earlier actions as real observations leave the rolling context. We introduce PARM-WM, a world model with Persistent Action–Response Memory (PARM). An observable response encoder combines visual feature changes with aligned actions to form a compact state. A zero-preserving, low-rank residual then reads this state throughout the rollout. A validation-selected latent interpolation combines the response branch with a JEPA continuation using one mixing weight across sources. We evaluate the complete predictor against seven action-conditioned baselines on three robotic trajectory sources, using seven observed frames to predict twenty future frames. PARM-WM achieves the best values on all four metrics for RoboMimic Lift, LIBERO New, and LanguageTable, reducing pixel MSE relative to JEPA-WM by 49.3%, 3.8%, and 12.1%, respectively. Globally Holm-corrected paired tests confirm improvements over every baseline on all four metrics on all three sources. These results establish persistent action–response memory as a useful interface for sustained visual forecasting.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.