acceptodds
Under review as a conference paper at ICLR 2027

W-DiT: Structure-Guided High-Frequency Propagation for Fashion Video Generation

Abstract

Generating fashion videos from reference images requires garment details to remain consistent under changing poses and cloth deformation. High-frequency product details, such as logos and fabric textures, should not be repeatedly regenerated across frames, as their appearance can progressively drift during motion. Our key insight is to generate low-frequency dynamics while propagating high-frequency product information directly from reference images along the garment's motion trajectories. Guided by this principle, we propose W-DiT, a video diffusion framework that couples generated garment motion with structure-guided detail propagation. First, frequency disentanglement separates low-frequency structure from high-frequency appearance and encodes reference details through a dedicated pathway that bypasses VAE compression. Second, structure-guided alignment matches low-frequency reference and target features to establish soft spatial correspondences that track the garment's evolving pose and deformation. Third, high-frequency propagation uses these correspondences to sample details directly from the fixed reference feature bank and inject them into each target frame, transporting product appearance along the motion trajectories. To support this task, we curate FashionVideo from authorized Xiaohongshu images, comprising approximately 10,000 synthesized fashion videos and 30,000 associated images with textual conditioning. Our experiments show that W-DiT preserves fine-grained garment details across changing poses while maintaining motion quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.