acceptodds
Under review as a conference paper at ICLR 2027

Refrain: Learning Multi-Shot Consistency from Inconsistent References for Video Generation

Abstract

A comic-book character, a photorealistic cityscape, an oil-painted coffee cup — such stylistically diverse references provide visual guidance across shots and frames, yet generating a coherent multi-shot video that unifies their visual styles while preserving subject identities and scene content remains challenging. Existing methods maintain cross-shot continuity through additional attention control or memory processing. However, they struggle to resolve disparate per-shot styles or handle reference–target duration mismatches within a single framework. To tackle this, we introduce **Refrain**, a unified model that learns cross-shot consistency through a self-supervised reconstruction task: reference clips from the same original video are deliberately perturbed with distinct visual edits, challenging the model to recover the unedited source footage under a shared style prompt. In this way, Refrain learns to suppress conflicting stylistic artifacts while faithfully retaining reference content and motion dynamics. Operating on the principle (**Any Reference, Any Duration, One Timeline)**, Refrain maps image and video references onto a shared temporal canvas without auxiliary conditioning modules. For evaluation, experiments show that standard CLIP or DINO metrics frequently overlook stylistic clashes or misinterpret intended narrative cuts (such as changes in subjects or scenes) as inconsistency. We therefore propose **Shot-Debiased Consistency (SDC)**, which uses an extensive real-video bank to establish a baseline for style variation associated with content changes within shots. By discounting this baseline divergence, SDC evaluates cross-shot style shifts against real-world distributions, where smaller residual differences indicate higher stylistic continuity. We further construct **DurMix-300K**, a dataset pairing stylistically altered clips of varying lengths with their original video sequences, to train the model. Extensive experiments demonstrate that Refrain achieves state-of-the-art performance in cross-shot consistency, while excelling at shot timing, reference adherence, and text alignment.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.