acceptodds
Under review as a conference paper at ICLR 2027

StoryWeaver: Towards Consistent and Synchronized Multi-Shot Audio-Video Generation

Abstract

Generating multi-shot videos with synchronized audio remains a significant challenge. While single-shot models perform well, extending them to multi-shot story generation is hindered by consistency degradation across shots. Current approaches typically rely on decoupled paradigms that generate segments independently, inherently limiting long-range coherence and causing identities, voices, and backgrounds to drift over time. Furthermore, existing continuous story frameworks often lack synchronized audio synthesis. To address these limitations, we introduce StoryWeaver, a framework for consistent multi-shot audio-video generation. Instead of decoupled steps, it adopts a joint generation paradigm conditioned on text and an initial prompt to improve global coherence. To model audio-visual interactions across extended contexts, we propose the Multi-Modal Sink Shot (MSS) mechanism, which couples windowed audio-visual cross-attention for local synchronization with sink-mediated inter-shot self-attention for cross-shot consistency. To extend generation beyond a single denoising window, we introduce Audio-Visual Latent Stitching (AVLS), a train-test matched continuation scheme that operates entirely in the joint audio-video latent space, reducing identity and acoustic drift compared to decode-reencode approaches. Experiments on both video-only and audio-visual benchmarks, supported by audio-specific ablations, adjacent-shot transition metrics, and human preference studies, show that StoryWeaver achieves improved visual-acoustic consistency, character preservation, and temporal coherence over existing methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.