Token Forcing: Geometry-Guided Search and Memory Selection for Autoregressive Video Generation
Abstract
Autoregressive video diffusion models generate long videos block by block, but over minute-long rollouts their scenes drift as the camera moves. Each block is committed once generated, so geometric errors accumulate, and what remains controllable is which content later blocks rely on. Existing approaches use the generator's own signals, require retraining, or score whole candidates with external rewards. Yet a generated block is rarely uniformly right or wrong, so the question is not which candidate to keep, but which parts of a single generation later blocks should build on. We introduce Token Forcing, a training-free method that answers this question token by token. During the early denoising stages of each block, a frozen recurrent 3D foundation model, whose state summarizes the committed history, keeps the tokens it judges consistent with the scene so far. The generator re-encodes these tokens at the matching noise level and injects their key–value features into the next block at the same denoising stage, where they act as soft geometric anchors. Because the 3D model only decides which tokens to use, neither model is trained. To measure drift over long camera trajectories, we introduce SceneLongBench, 100 static indoor and outdoor scenes for 60-second text-to-video and text-and-image-to-video generation. Relative to Causal Forcing, Token Forcing preserves video quality on motion-rich prompts while reducing reprojection error in all four SceneLongBench settings and epipolar error by up to 21.6%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.