acceptodds
Under review as a conference paper at ICLR 2027

FlowVVTON: Flow-Guided Mask-Free Video Virtual Try-On

Abstract

Video virtual try-on aims to transfer a target garment onto a moving person across video frames. Existing methods often rely on human parsing masks or pose keypoints that can become unreliable under large motions and occlusions, causing boundary artifacts and temporal inconsistency. A further limitation is that most approaches rely solely on attention mechanisms for temporal modeling, providing no explicit motion supervision. To address both limitations, we propose FlowVVTON, a mask-free framework that eliminates parsing mask dependency while introducing explicit motion supervision. Our key innovation is a flow-warped latent loss that uses optical flow solely during training to align features across frames at multiple layers of the generation model, enforcing multi-scale temporal consistency. We incorporate this loss into a two-stage training strategy that first establishes mask-free spatial alignment and then introduces flow-guided temporal supervision. Together, these designs improve temporal consistency and reduce boundary artifacts, enabling robust garment transfer under large body motions. Experiments on TikTokDress show that FlowVVTON outperforms baselines by substantial margins, including a 5.7 reduction in unpaired VFID-R compared with SwiftTry, while requiring no segmentation masks, pose keypoints, or region annotations as model inputs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.