ReDreamer: Geometry-Grounded Memory for Real-Time Driving World Models
Abstract
Video world models should generate evolving scenes in real time while preserving object appearance over time. Yet objects can reappear with altered appearances despite correct motion and geometry. The model reproduces where they are, but not what they looked like. Our key idea is to use the persistent object identities and geometry in scene controls as addresses for appearance memory. ReDreamer is a chunk-wise autoregressive video diffusion model that records generated appearances and retrieves them at current object locations through geometry-grounded attention. However, training with real-image references while reading generated appearances at inference introduces exposure bias in memory conditioning. We address this mismatch through rollouts that build and read their own memory, using distribution matching distillation to supervise continuations without forcing generated appearances to match the original dataset video. On our nuScenes-derived re-entry evaluation, ReDreamer improves our Re-entry Appearance Consistency (RAC) score by 10.5% over the strongest evaluated baseline. Enabling memory reads improves RAC by 12.0%. With four denoising steps per chunk, it reduces FVD by 5.7% relative to MagicDrive-V2 in our front-view evaluation and generates at 8.1 FPS.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.