ReWorld: An Interactive World Model with Long-Horizon Memory
Abstract
Interactive world models must follow actions, recall past views, and stream in real time. Control needs a short attention horizon; memory needs a long one. Jointly learning these abilities can strengthen action following while weakening revisit consistency. ReWorld separates these during training through mixed per-head attention windows and random head routing. Routing exposes each head to both horizons, allowing all heads to share a bounded cache at deployment without additional model parameters. Chunk-drop training prepares the model for sparse histories, while a bounded KV cache and pose-indexed landmark bank support long-horizon recall at inference. The bank preserves selected past chunks and retrieves them by camera pose, retaining spatially relevant history beyond the recent window. A metric-aligned data pipeline unifies eight synthetic, game, and real-world sources, and LoRA-confined distillation enables both high-fidelity and real-time 704 × 1280 generation from one backbone. Against six recent interactive world models, ReWorld achieves the lowest rotation error (11.95◦), best camera-motion consistency, and highest mean score across the seven evaluated VBench dimensions. On 64 s out-and-back rollouts (384 latents), its fixed 12-chunk cache still regenerates the starting view.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.