DualFashion: Dual-Sided Garment-Conditioned Multi-View Fashion Model Generation via Hierarchical Structured Inpainting
Abstract
Most virtual try-on methods edit a supplied full-body person image, while garment-driven virtual dressing typically generates a single image from a frontal garment reference. Neither setting directly supports creating a coherent multi-view fashion presentation from complementary garment evidence. We introduce DualFashion, a framework for dual-sided garment-conditioned multi-view virtual dressing. Given paired front and back garment photographs and a head reference, DualFashion directly generates an ordered set of natural-presentation, front, side, and back views, without requiring a target body image, pose, or scene. The task requires accurate routing of front- and back-side garment details, consistent garment and subject appearance across views, and sufficient spatial capacity for fine-grained details. We address these challenges through hierarchical structured inpainting. A shared-canvas joint stage first synthesizes a coordinated four-view set within a single denoising process. A set-conditioned view-wise stage then reconstructs each view at higher resolution using the complete intermediate set together with the original garment and head references, improving local details while preserving set-level coherence. To support this task, we introduce DVF-6K, a dataset of 6,011 product groups across 33 categories, pairing front and back retail garment photographs with natural-presentation, front, side, and back on-body views. Experiments demonstrate improved visual quality, garment fidelity, identity preservation, viewpoint correctness, and cross-view consistency. The generated views can also serve as sparse conditions for downstream fashion-video synthesis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.