SceneGen: Generating Multi-Shot Scenes with Timed Speech and Sound
Abstract
Recent advances in joint audio-video generation have enabled realistic scenes with synchronized speech and sound. However, fine-grained temporal control remains challenging for open models, especially when camera transitions, visual actions, and audio events exhibit distinct temporal patterns within the same scene. Global conditioning cannot define local timing, while shot-level conditioning restricts events to camera boundaries, disrupting cross-shot continuity. To address this limitation, we propose \MethodName, a framework for fine-grained temporal conditioning that models shot transitions and audio events separately. Event-Aware Attention independently encodes complete speech and sound events, activating them based on their timeline intervals. Shot-Aware RoPE explicitly represents camera transitions through shared temporal position shifts across the audio and visual modalities. To provide supervision consistent with these mechanisms, we construct \DatasetName, containing approximately 194K multi-shot scenes with raw audio, linked character identities across views, shot-level visual annotations, and independently timed audio events. Together, these designs allow the model to follow shot-specific instructions while preserving events across camera cuts. Experiments on \BenchmarkName demonstrate improved speech timing, prompt following, and cross-shot character consistency, maintaining overall audiovisual quality. We will make our dataset and models publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.