FACT: FREQUENCY AND ATTENTION CONTROL FOR TRAINING-FREE VIDEO STYLIZATION
Abstract
Reference-guided video stylization transfers the appearance of a reference image to a source video while preserving source structure and temporal consistency. Recent training-free approaches leverage pretrained diffusion models to avoid costly training, yet they can lose fine-grained source details during style transfer and produce inconsistent appearance across frames. We analyze the frequency components of intermediate denoising latents and find that low-frequency components evolve gradually while high-frequency components change more rapidly across steps. Motivated by this finding, we propose FACT, a training-free framework that combines latent frequency modulation with attention control. At each denoising step, frequency modulation attenuates the low-frequency component of the current latent while retaining its high-frequency component, preserving fine spatial details throughout stylization. Logarithmic Length Attention Scaling (LAS) adaptively adjusts content-to-style attention strength across U-Net feature scales, balancing style injection with source-structure preservation. Confidence-Weighted Temporal Attention (CTA) guides cross-frame attention along tracked point correspondences, weighting interactions by tracking confidence and temporal distance to improve frame-to-frame consistency. Experiments show that FACT outperforms existing methods in reference-style transfer, source-structure preservation, and temporal consistency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.