acceptodds
Under review as a conference paper at ICLR 2027

Closing the Loop: Training-Free Revisit Consistency for Autoregressive Generative Rendering

Abstract

Recent conditional video generation models have shown promising potential to transform 3D engine renderings, such as depth maps and untextured geometry, into photorealistic videos for diverse gaming and immersive content creation. These applications require long-horizon autoregressive generation that continuously synthesizes new frames while preserving a persistent 3D world. Autoregressive generators synthesize video chunk by chunk with a bounded KV cache, so when the camera revisits a location after its context has been evicted from the KV cache, the model often regenerates inconsistent appearance, even though the conditioning renderings (e.g., depth) remain perfectly aligned since they originate from the same underlying 3D geometry. We address this revisit inconsistency without any post-training by exploiting correspondences the 3D engine already provides: temporal correspondence retrieves pose-matched historical latent chunks into the KV cache as loop-closure memory, while spatial correspondence from camera pose and depth reprojection biases token-level attention toward geometrically corresponding regions of the retrieved chunks. We demonstrate our method on loop-closure trajectories mined from the TartanAir and TartanGround datasets to mirror complex real-world application scenarios, where it outperforms existing training-free baselines on revisit consistency without losing video quality.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.