Geometry Tells Where; History Tells What: Appearance Memory in Causal Video Generation
Abstract
Video generators now demonstrate impressive capabilities in creating controlled qualitative content. However, camera-controlled video generators still produce substantial appearance drift when revisiting earlier viewpoints. In this paper, we build a competitive geometry- and camera-conditioned causal backbone and use it to investigate what information should be preserved and where recalled memory should enter the generator. We find that geometry alone is insufficient, even when it covers almost the entire view: appearance is recovered reliably only when previously generated RGB image is recalled and temporally aligned with the frame currently being generated. Crucially, simply providing the correct recalled RGB is not enough: when its temporal position is reset to the beginning of the sequence, rather than current ,its benefit largely disappears, performing no better than having no memory. Based on these findings, we introduce CHaRM, a training-free appearance memory that stores generated RGB frames indexed by camera pose. At revisit time, CHaRM reuses the stored frames when the camera returns to a stored pose, or otherwise warps nearby stored frames into the new view, gives them the temporal position of the frame being generated, and injects them next to the KV cache of recent frames. As a result, CHaRM cuts self-revisit LPIPS (consistency with previously generated views) from 0.504 to 0.099 on exact revisits, and brings return LPIPS (fidelity to ground truth on the way back) to the level the backbone reaches in a single uninterrupted pass, i.e., with no loss of quality from revisiting.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.