acceptodds
Under review as a conference paper at ICLR 2027

CAVE: A Generated World Should Still Explain Its Past

Abstract

Video generation models can produce visually realistic sequences, yet object structures and spatial relationships may still drift over time. Existing geometric metrics compare frame pairs or jointly model the entire sequence under a shared interpretation, making conflicts difficult to localize or allowing anomalous observations to be absorbed into the overall solution. We introduce an observer-centric principle: video geometric consistency should be evaluated in observation order. The historical prefix provides an internal geometric reference, and a new observation should preserve geometric relationships already supported by history. Building on this principle, we develop CAVE, which maintains an evolving world state and adopts an “assimilate-then-look-back” verification strategy. At each timestep, CAVE queries identical historical views before and after the update and measures the resulting increase in historical explanation error. These responses form a geometric damage tensor that supports video scoring, temporal conflict localization, and current-view spatial attribution. On controlled geometric conflicts, CAVE improves temporal and spatial localization over prior metrics, with particularly strong gains under persistent conflicts. We further use the CAVE video score as preference supervision for DPO post-training, improving geometric consistency across multiple independent metrics while maintaining video quality and dynamics. CAVE therefore turns directional historical verification into a practical signal for both geometric diagnosis and post-training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.