Looking Below the Latent Token: Detail Retrieval via Correspondence Supervision for Virtual Try-On
Abstract
Latent-space virtual try-on (VTON) models often distort fine-grained garment details, such as small text, logos, stitching, and repetitive patterns. While existing methods improve garment-person correspondence, such correspondence is defined over spatially coarse latent tokens and therefore provides only coarse localization of reference garment content. We introduce DECOR-VTON, a two-phase framework that uses latent correspondence as a spatial routing signal for selectively accessing finer garment information. Through layer-wise analysis and controlled attention interventions, we identify late dual-stream blocks as promising correspondence-aware representations. Phase I improves correspondence in these blocks using probabilistic warp-consistency supervision without requiring dense correspondence annotations or an external correspondence teacher during training. In Phase II, the refined correspondence guides the selection of relevant garment regions, from which higher-resolution convolutional features are retrieved through local cross-attention and injected into downstream transformer blocks via a timestep-conditioned residual gate. This coarse-to-fine design complements latent-space generation with selectively accessed fine-grained reference features without introducing global high-resolution attention. Experiments on VITON-HD and DressCode show consistent improvements in overall synthesis quality, while qualitative comparisons indicate better preservation of fine-grained garment structures. Extensive ablations further demonstrate the complementary roles of correspondence supervision and high-resolution detail retrieval.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.