Reveris: Adaptive Reference Conditioning for Few-Step Visual Effect Transfer
Abstract
Advances in visual generative models have expanded digital content creation, where creators seek to reuse distinctive effects across subjects and scenes. Reference videos provide concrete demonstrations of transformations that are difficult to specify through text alone, but also entangle the intended effect with reference-specific content. Faithful visual effect transfer therefore requires capturing the relevant change evidence while adapting its influence to the target. We introduce Reveris, a four-stage training framework spanning reference adaptation, modulation calibration, multi-source guidance, and few-step distillation. Within a shared image-to-video backbone, we combine clean-reference temporal variation with caption-conditioned content matching and target-conditioned modulation to adapt reference evidence to new content. We further coordinate target and reference conditions through multi-source classifier-free guidance, and distill this guided behavior through distribution matching into a four-step generator. We also construct a curated dataset of 49,194 videos spanning 350 effect types in 13 families to support learning shared effect behavior across diverse subjects and scenes. Experiments demonstrate strong overall transfer performance, with consistent improvements in effect fidelity, source preservation, and visual quality. Our project demo is available at https://reveris123.github.io/reveris.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.