Cine-Long: Consistent Long-Horizon Multi-Shot Audio-Visual Generation
Abstract
Long-horizon, multi-shot audio-visual generation remains challenging in both data and modeling. Although modern generators can already produce multiple shots within a single generation chunk, existing datasets are not explicitly organized around such chunks or the cross-chunk dependencies encountered during iterative long-video generation. On the modeling side, sustaining a coherent long video requires selecting the relevant and reliable visual and audio context from an ever-growing generation history. We address both challenges with three contributions. First, we introduce a long-video dataset annotated at the chunk level, where each chunk spans multiple shots, aligning supervision with how modern models generate and targeting cross-chunk coherence. Second, we present Cine-Long, which equips the generator with an audio-visual memory bank as explicit context during training, so it learns to condition on key frames and audio drawn from past chunks rather than on a fixed window. Third, at inference we introduce a memory harness, an agentic controller that steers each target generation by retrieving, from the model's own generated history, the key frames and audio segments most relevant to the current shot—under a fixed memory capacity and without rewriting the script or re-sampling the generator. To evaluate this setting, we build CineLong-Bench, a benchmark for multi-shot audio-visual generation with character re-entry and voice continuation. Cine-Long achieves state-of-the-art coherence and audio-visual consistency on CineLong-Bench.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.