FineStyler: Towards Detail-Preserving Video Stylization
Abstract
Reference-based video stylization transfers the style of a reference image to an input video while preserving its content and motion. Existing video diffusion methods often rely on overly compressed style representations and coarse structural conditioning, leading to washed-out micro-style details, reference content leakage, and temporal drift of fine structures. We present FineStyler, a single-stage inference framework for high-fidelity, temporally stable video stylization with a single style reference image. First, Content-Anchored Style Modulation aligns learned style embeddings to a natural-image feature domain, encouraging fine-grained texture encoding while suppressing instance-specific content leakage. Second, High-Frequency Adapted LoRA injects input-video high-frequency cues into DiT self-attention through a parameter-efficient adapter, preserving thin edges and local geometry beyond silhouette-level controls. Third, a style reference image curation pipeline synthesizes style-aligned yet content-mismatched references from paired raw and stylized videos using image stylization models, then selects high-quality references through automatic scoring to enable training that matches the inference interface. Extensive experiments demonstrate that FineStyler improves style faithfulness, content preservation, and temporal stability over prior diffusion-based baselines, producing sharper structures and higher perceptual quality in challenging content-mismatched reference settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.