SartorAI: Multimodal Garment Pattern Generation and Editing with Tokenized Geometry
Abstract
Garments and their sewing patterns play an important role in fashion, gaming, film, and virtual reality. However, designing garments that can be sewn together still requires considerable expertise, making it difficult to meet the growing demand for personalized clothing. To bridge this gap, we propose SartorAI, a framework for garment generation and editing. Rather than representing garment panels as sets of parametric edges, we model their boundaries as arbitrary closed 2D curves and define stitching relationships as connections between boundary points. This representation avoids dependence on predefined panel-edge parameterizations and stitching schemes, which are commonly used in existing works. Our key insight is that panel shapes and their 3D layout provide strong cues about how the panels should be sewn together. We therefore separate panel geometry generation from stitching prediction, using a vision-language model (VLM) to generate panel geometry and a flow model to predict stitching conditioned on that geometry. To improve geometric precision, we encode panel geometry into discrete tokens using residual vector-quantized variational autoencoders (RQ-VAEs) and train the VLM to predict these tokens. This preserves the VLM's symbolic reasoning capabilities and interactivity while alleviating the precision limitations of autoregressive floating-point prediction. To train SartorAI, we construct a dataset of 100k garments, each draped over human bodies with diverse shapes and rendered with realistic PBR textures. Trained on this dataset, SartorAI achieves state-of-the-art performance in garment sewing pattern generation and editing.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.