acceptodds
Under review as a conference paper at ICLR 2027

TTT2Reel: End-to-End Test-Time Training for Long Video Generation

Abstract

Generating a long, coherent video requires preserving visual information about previously-generated frames in the window beyond a model's bounded context window. We present TTT2Reel, an end-to-end test-time training framework for chunk-wise autoregressive long video generation that stores evolving visual history in lightweight LoRA parameters. Before generating each new chunk of a video, the model updates these parameters by optimizing a flow-matching loss on the preceding chunk, while keeping the pretrained backbone frozen. We train the model through bilevel optimization, where an inner optimization loop updates the parametric memory through test-time training, and the outer loop optimizes its initial parameters to minimize subsequent prediction losses across the video. To withstand compounding errors in self-generated context at inference, we estimate model-dependent residuals online and replay them into both the memorization and prediction steps during training. Experiments on minute-long video generation show that TTT2Reel achieves the best average rank across six VBench-Long dimensions among the evaluated methods, balancing dynamic content with visual quality and consistency.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.