Real-time Streaming Video Stylization via Rectified Style Flow
Abstract
Diffusion-based image stylization provides strong appearance control but is too slow and temporally unstable for streaming video. Directly training a causal video generator, on the other hand, entangles style, motion, and long-horizon autoregression and makes every new style expensive to support. We introduce RSF, a staged framework for real-time streaming video stylization. Its core is rectified style flow: rather than generating from pure noise, the model transports a partially corrupted source distribution at noise level s to a target-style distribution. We first learn modular style LoRAs on paired source-target video clips. We then freeze a bank of these LoRAs and train the full video DiT with teacher ODE pairs and causal forcing, separating reusable autoregressive dynamics from style-specific parameters. The resulting causal teacher is distilled to a one-step student using distribution matching distillation, followed by generated-history rollout fine-tuning to reduce exposure bias. Finally, a lightweight recovery procedure adapts previously unseen style LoRAs without repeating the complete video training pipeline. At 832x480 resolution, optimized mixed-FP8 inference sustains 102.1 FPS with 78.4 ms input-to-output latency on a single NVIDIA H100 and 36.3 FPS with 220 ms latency on a single NVIDIA RTX 5090. The H100 deployment requires only 4.81 GiB of steady-state resident GPU memory. We evaluate target-style text alignment, source preservation, and temporal stability on a 50-video benchmark under five styles against four public video stylization systems; the strictly paired CLIP-T comparison uses 47 complete videos. We will open-source all training code, model weights, and prepared datasets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.