CastDirector: Directing Joint Audio-Video Generation with Grounded Screenplays
Abstract
Cinematic audio-video generation must assign dialogue and actions to the intended characters while producing immersive, scene-consistent audio. Existing systems can produce coherent clips while assigning speech or actions to the wrong character or omitting requested events. We introduce CastDirector, a screenplay-driven system for coordinating character performances and cinematic soundtracks. A grounded screenplay specifies the cast, speaker identities, dialogue, actions, and camera and audio cues as ordered events. To construct training screenplays from audiovisual data, our annotation pipeline extracts transcripts and visual actions independently, then verifies and combines them with the original clip to bind dialogue and actions to characters. The proposed multimodal Director jointly encodes this screenplay and the first frame, conditioning both generation streams to connect character actions, voices, and sound events to the visible scene. We first align a pretrained LTX-2.3 generator with the Director's features, then adapt the Director with LoRA through the audio-video generation loss. Under matched screenplay content, experiments show improved speaker binding, higher human ratings for action, dialogue, and audio, and gains in audio quality and speech accuracy. Qualitative examples showcase coordinated performances and immersive audio combining character voices, environmental ambience, sound effects, and scene-appropriate music.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.