RaAVGen: Rollout-Aware Long-Form Audio-Video Narrative Generation from Global Scripts
Abstract
Generating long-form audio-video narratives from a complete script requires tracking what has already been generated and continuing unfinished content across generation boundaries. Segment-specific prompting workflows assign this responsibility to separately prepared local prompts. We introduce a rollout-aware global script framework that jointly generates audio and video segment by segment while retaining the complete script throughout generation. Our key insight is that visual continuity and script localization require different temporal contexts. We therefore pair a short video history that captures the recent visual state with a longer audio history that helps the model locate preceding speech in the script and determine what to generate next. This asymmetric design preserves extended spoken context while limiting the token cost of dense video history. For multi-shot narratives, sparse keyframes from completed shots provide earlier scene and character appearance, while relative shot timestamps guide shot transitions. Together, these conditions allow generation boundaries to fall within shots, utterances, and actions. To address limited long-form training data, we combine complementary supervision with staged training to integrate face and voice conditioning, shot transitions, and long-sequence continuation. We further construct pseudo-long scripts by extending speech text around existing captions while retaining the original audio-video supervision, training script localization without matching full-length recordings. Experiments demonstrate long-form multi-shot generation from a complete script without separately prepared local prompts, alongside support for optional face and voice personalization in single-person and two-person settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.