acceptodds
Under review as a conference paper at ICLR 2027

CineRef: Cinematic Reference-Guided Audio-Video Generation

Abstract

Cinematic reference-guided audio-video generation aims to synthesize high-quality videos conditioned on reference images and audios under complex cinematic scenarios. However, current open-source methods primarily treat reference information as passive conditioning, making reference injection fragile under large camera motions and frequent shot transitions. To address this, we propose CineRef, a unified framework for cinematic reference-guided audio-video generation. The key insight is that reference information should be explicitly linked to the generated video during training and consistently preserved throughout the entire clip, rather than treated as passive and transient conditioning. Based on this insight, we first propose Dual-Anchor Reference Alignment, which combines a reference ID anchor for explicit ID supervision with a face-aware content anchor for preserving natural facial content. We then introduce Reference Trajectory Consistency to maintain long-range ID coherence across substantial viewpoint and scale variations along ID-wise trajectory. Finally, we design Audio-Visual Reference Binding to establish explicit correspondence between each visual ID and its corresponding audio reference. Extensive experiments on our proposed CineSet for audio-video generation demonstrate that CineRef achieves superior ID injection and preservation under diverse camera motions and shot transitions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.