EverStyle: Long-Horizon Consistent Streaming Video Stylization
Abstract
Streaming video stylization continuously transforms input video frames into user-specified visual styles, supporting real-time applications such as live broadcasting and interactive content creation. Recent diffusion-based methods have made substantial progress toward efficient and temporally coherent streaming stylization, yet maintaining consistent stylization throughout an open-ended rollout remains challenging. In particular, previously seen content is usually stylized differently when it reappears after leaving the active context, despite smooth local segments. To address these, we propose **EverStyle**, a novel long-horizon consistent streaming video stylization framework. To recover the established appearance of recurring content, we design Source-Indexed Stylization Memory (SISM), which stores and retrieves source-indexed stylization residuals, and Replay-Based Stylization Consistency Regularization (RSCR), which trains the generator to reproduce earlier stylization under different temporal contexts. Yet reproducing earlier appearance alone, without constraining new content to the same distribution, cannot keep the evolving stylization temporally consistent, leading to global drift. To suppress this, we introduce Early-Anchored Stylization Distribution Matching (EA-SDM), which jointly evaluates early and current generations to align with the same style without content correspondence. Experiments demonstrate that EverStyle maintains competitive stylization quality, while substantially improving long-horizon consistency under both natural scene revisitation and exact content recurrence settings, sustaining consistent stylization over sequences exceeding 130,000 frames.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.