BiCorr: Learning Bidirectional Latent Correspondences for Video Virtual Try-On
Abstract
Video virtual try-on transfers a reference garment to a person in a video while preserving identity, motion, background, and garment appearance. Recent video diffusion Transformers have improved garment conditioning and temporal modeling, but fine-grained patterns and local garment structures can still be distorted, particularly under large motion and occlusion. Garment conditioning provides reference information but does not ensure that each person token retrieves relevant features from the correct garment region. To address this issue, we propose BiCorr, a framework for bidirectional latent correspondence learning between person-video and reference-garment tokens. We combine high-confidence direct matches with matches propagated along validated trajectories, and balance their coverage across time and space. These matches are mapped to the model's video and garment tokens and represented as distributional targets that account for mapping uncertainty while retaining distinct garment locations. The targets guide person tokens to retrieve features from corresponding garment regions, while a reverse objective associates garment tokens with their corresponding spatiotemporal locations throughout the video. All matching and correspondence losses are used only during training, requiring no additional matcher, tracker, or warping module at inference. Experiments show that BiCorr improves visual quality, garment fidelity, and temporal consistency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.