acceptodds
Under review as a conference paper at ICLR 2027

FACT: FREQUENCY AND ATTENTION CONTROL FOR TRAINING-FREE VIDEO STYLIZATION

Abstract

Reference-guided video stylization transfers the appearance of a reference image to a source video while preserving source structure and temporal consistency. Recent training-free approaches leverage pretrained diffusion models to avoid costly training, yet they can lose fine-grained source details during style transfer and produce inconsistent appearance across frames. We analyze the frequency components of intermediate denoising latents and find that low-frequency components evolve gradually while high-frequency components change more rapidly across steps. Motivated by this finding, we propose FACT, a training-free framework that combines latent frequency modulation with attention control. At each denoising step, frequency modulation attenuates the low-frequency component of the current latent while retaining its high-frequency component, preserving fine spatial details throughout stylization. Logarithmic Length Attention Scaling (LAS) adaptively adjusts content-to-style attention strength across U-Net feature scales, balancing style injection with source-structure preservation. Confidence-Weighted Temporal Attention (CTA) guides cross-frame attention along tracked point correspondences, weighting interactions by tracking confidence and temporal distance to improve frame-to-frame consistency. Experiments show that FACT outperforms existing methods in reference-style transfer, source-structure preservation, and temporal consistency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.