CodingImage: Programmable Image Editing via Vector Graphics Code
Abstract
Fine-grained image editing calls for intuitive, element-level control. Scalable Vector Graphics (SVG) code offers a shared editing interface: large language models (LLMs) can rewrite it from natural-language instructions, while humans can directly refine its paths. However, conventional image-to-SVG methods face two challenges: their outputs may lack the semantic alignment and hierarchy needed for intuitive editing and manipulation, and pursuing high reconstruction fidelity can make vectorization prohibitively time-consuming. We introduce CodingImage to address both challenges. It efficiently parses images into hierarchical vector graphics (VGs) that are semantically aligned and structurally coherent. For photorealistic synthesis, we further propose Vector-Guided Noise Prediction, which conditions the diffusion model on these VGs and guides generated images to follow their structure and semantics. This enables precise control over geometry, color, and object semantics. Experiments demonstrate CodingImage's effectiveness in image editing, object-level manipulation, and fine-grained content creation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.