Anamnesis: Training-Free Episodic Recollection for Autoregressive World Models
Abstract
Autoregressive (AR) world models generate interactive worlds by conditioning on the recent past, yet they systematically forget it: scenes revisited after minutes of exploration are regenerated as if seen for the first time, and spatial memory collapses within the span of a single attention window. We argue that this failure has a precise phenomenological signature: current world models implement retention (the KV cache) and protention (action conditioning), but possess no recollection: the deliberate reactivation of earlier experience when the present re-encounters it. We introduce , a training-free episodic recollection mechanism for frozen AR world models: a buffer stores denoised latent states tagged with exact camera poses; a pose-gated detector recognizes revisits; and recognized revisits are re-rendered by re-noising the stored latent to a low level and granting the model a single refinement step. No gradients are updated. A KV-cache alternative (writing the anchor into the recency slot) leaves the output unchanged: recollection must enter through the generative channel, not the retention channel. To measure recollection we propose a closed-loop revisit benchmark with analytically exact pose ground truth from the model's own action interface, and an Excess Revisit Consistency metric that controls for the slow-motion shortcut. On the open-source LingBot-World model (one 4-GPU node), the baseline's excess is negligible (mean ERC-PSNR dB; negative on the pan loop) while lifts it to dB on all four loops (), improving PSNR, LPIPS, and SSIM throughout, with a random-anchor control showing the effect is carried by pose-gated detection. Widening the KV window up to does not restore what the excursion erased: recollection, not longer retention, is the missing capability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.