DesignX: Bridging Coding and Generation for Recommendation-based Visual Design
Abstract
Despite recent advances in AI visual generation, professional design still faces two challenges: enabling non-experts to create professional-quality designs and converting raster images into unified, editable representations of raster graphics, vector graphics, and text. To address these challenges, we present DesignX, a recommendation-based visual design agent that enables both professional design creation and editable visual reconstruction. For the design agent, we construct a large-scale template library covering diverse visual design categories and develop three specialized skills: a cataloging skill for organizing and continuously expanding the template library, a coarse-to-fine retrieval skill for efficiently identifying relevant templates, and a generation skill that adapts templates according to user requirements. Beyond raster-only generation, we introduce an editable visual decomposition framework consisting of a vision-language model (VLM), an RGBA-based diffusion transformer (DiT), and a text-attribute prediction network. Specifically, we design a structured decomposition protocol that enables the VLM to parse any design into seven editable elements with spatial relationships and semantic descriptions while directly generating code for simple vector primitives. The RGBA-based DiT extracts individual layers from VLM descriptions and completes occluded regions, while the text-attribute prediction network jointly predicts typography attributes through a multi-head Transformer architecture. Furthermore, we introduce synthetic collage data augmentation to improve model robustness. Finally, we establish a comprehensive visual design benchmark spanning diverse categories to facilitate future research. Extensive experiments demonstrate that DesignX significantly improves visual design quality, and our decomposition framework surpasses existing SOTA commercial solutions. We have released the code and template library to promote research in this direction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.