SymFashion: Learning Cross-Task Symmetry for Unified Fashion Generation
Abstract
Can compatible forward and reverse fashion transformations provide useful supervision for one another, or does joint training merely share parameters? We address this question by proposing SymFashion, a unified fashion-generation model built on a single rectified-flow transformer that supports cloth-to-person try-on (C2P), person-to-person (P2P) garment transfer, virtual Try-Off, and pose-controllable cloth-to-model (C2M) generation. SymFashion maps compatible directions to a shared human–garment state and explicitly couples their role-aligned clean endpoints at the same timestep using permutation-aligned noise. Accounting for the latent mismatch introduced by the two encoding paths, our analysis recovers -weighted velocity agreement under exact latent alignment and supplies a descriptive cross-view root-risk bound. On VITON-HD and DressCode, a single checkpoint remains competitive across all four task families. Under a fixed total optimization budget, uncoupled joint training has mixed effects on the primary task-level readouts: it degrades C2P and C2M but improves P2P and Try-Off, whereas symmetric coupling improves both regimes. Direct diagnostics reveal substantially lower cross-view disagreement, especially at higher noise levels and in person-side garment regions, despite higher branchwise prediction risk. Thus, compatible transformations can support one another, but the benefit is more consistent across our evaluations when their shared semantic structure is explicitly coupled rather than left to uncoupled joint training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.