CineAstra: References as In-Context Directors for Multi-Shot Video Creation
Abstract
Cinematic multi-shot video generation requires both holistic control over cinematic language and versatile support for diverse creation modes. Existing methods largely rely on textual prompts or condition-specific modules, making it difficult to convey complex directorial intent and flexibly accommodate partially observed footage. To address these limitations, we introduce **CineAstra**, a unified framework for reference-guided multi-shot video creation. **CineAstra** leverages in-context learning through a Mixture-of-Transformers (MoT) architecture, where reference and target video streams interact via contextual attention to infer the reference's cinematography and narrative pacing for the target screenplay. To further support flexible creation, we propose frame-wise flow timestep conditioning, which independently manages per-frame noise levels according to assigned diffusion timesteps, seamlessly integrating target observations at arbitrary temporal positions. Extensive experiments demonstrate that **CineAstra** outperforms existing methods in cinematography adherence, narrative pacing alignment and visual quality, providing a unified solution for versatile multi-shot creation modes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.