FusedVTON: Early Image Fusion for Efficient Virtual Try-On
Abstract
Existing image-based virtual try-on methods typically use explicit garment warping or model garment–person correspondence through reference networks, cross-attention, or self-attention over concatenated tokens. Although effective, these designs introduce architectural complexity and computational overhead, especially in multi-garment try-on, where additional garments require additional warping operations or garment representations. We view garment transfer as coarse spatial placement followed by non-rigid generative refinement and propose to implement this formulation. Reference garments are resized and placed within the clothing-agnostic person image before VAE encoding. The resulting composite is processed through a single, fixed-length image-token stream without a separate garment branch or appended garment tokens. Experiments on VITON-HD, DressCode, DressCode-MR, and Garments2Look show competitive quality. Separate profiling at without classifier-free guidance measures 53.9% and 51.7% fewer model-inference FLOPs than the corresponding SD1.5 and FLUX.1 CatVTON variants. At , counted model-inference computation remains constant across one to four garment inputs in the evaluated scaling setup.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.