Dressing the Crowd: Multi-Human Virtual Try-On via Low-Substitutability Preference Optimization
Abstract
Virtual try-on (VTON) has achieved impressive fidelity in single-person settings, but extending it to multi-human scenes requires resolving person–garment assignments, inter-person occlusions, spatial interactions, and non-target preservation, for which paired supervision is difficult to scale. We formulate Multi-Human VTON and introduce MultiHuman-VTON, a 55K-image dataset and benchmark covering one-to-many Partial and many-to-many Full editing. We further propose CES-VTON, a low-substitutability preference optimization framework that post-trains an image-editing foundation model using group-relative reward signals. Conventional additive rewards permit a strong criterion, such as source preservation, to compensate for failed editing, enabling reward-hacking solutions. To restrict such compensation, our Hierarchical CES Utility (HCU) applies nested CES aggregation: it first combines source-relative binding gain with source-garment replacement and then couples the editing utility with non-target preservation. Dynamic Fréchet Credibility (DFC) maintains synchronized source and generated feature queues to estimate each candidate's marginal effect on the source–generation feature-distribution gap, asymmetrically penalizing candidates that increase it. Experiments show CES-VTON outperforms specialized and general-purpose baselines in complex multi-human scenes, improving target assignment, garment replacement, and non-target preservation while its components generalize to single-person VTON and multi-image editing.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.