SEGUE: Visual Signal as Exclusive Guidance for One-Shot Transition Effect Generation
Abstract
One-Shot Transition Effect Generation (TEG) aims to bridge two video clips using a transition effect provided in a reference clip. However, existing models learn a text shortcut: they rely heavily on prompt descriptions, causing the visual transition signal to vanish. We present SEGUE, a framework designed to read the transition strictly from the visual demonstration via two key mechanisms: i) Visual reference isolation during training, pairing inputs with effect-free neutral prompts and applying a one-way attention mask that keeps the reference independent of the target; and ii) Null-reference guidance at inference, which defines the unconditional path as a dissolve of the reference endpoints—adding no motion or material of its own—to isolate and amplify the visual transition effect. SEGUE flexibly supports endpoints that are video clips, static images, or absent, unifying two-sided TEG and single-sided visual effect transfer within a single model. We also introduce the SEGUE-Benchmark, a cross-paired dataset of 1,692 transition classes and 56,368 training pairs spanning creative presets, parametric procedural transitions, and visual effects. On held-out effects, SEGUE outperforms baselines adapted to TEG and prior visual effect transfer methods even without the transition prompt, serving as the first open-source framework for one-shot TEG.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.