Seeing More Details: Multi-Reference Virtual Try-On with Fine-Grained Garment Representation
Abstract
Virtual try-on (VTON) has achieved remarkable progress with the advancement of diffusion models. However, most existing methods focus on single-image guidance under in-shop settings, often producing unsatisfactory results when handling large pose variations, complex occlusions, or fine-grained texture transfer. To address this limitation, we propose a mask-free VTON method, termed MR-VTON, which leverages Multi-Reference garment images captured from diverse viewpoints and wearing poses to preserve the garment details and realistic try-on appearance, particularly in in-the-wild scenarios. The proposed model employs a dual-branch architecture to dynamically adapt reference image selection and fine-grained feature aggregation. Specifically, the first branch introduces a router-based garment selection mechanism that adaptively selects a single image as the primary reference to anchor garment identity features, with minimal occlusion and maximum clothing information. The second branch introduces a target pose-guided feature aggregation module that integrates garment details and viewpoint information from multiple reference images, thereby enhancing spatial consistency and fine-grained detail preservation. The features from both branches are concatenated and injected into a Diffusion Transformer architecture as reference features, facilitating VTON across diverse viewpoints and poses. Furthermore, we construct an evaluation benchmark for in-the-wild multi-reference VTON, named MRG-4K, to fill the gap in this emerging task. Extensive experiments on the proposed benchmark and zero-shot academic benchmarks demonstrate that our method consistently outperforms existing approaches and achieves state-of-the-art performance in both single-reference and multi-reference try-on tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.