acceptodds
Under review as a conference paper at ICLR 2027

VA-SVG: Improving SVG Code Prediction via Generative Visual Alignment

Abstract

Existing SVG generation methods typically adopt multimodal large language models (MLLMs) as their backbone and formulate SVG synthesis primarily as code sequence prediction. However, SVGs constitute a distinctive image–code multimodal representation, requiring the model to generate code that is visually consistent with the target image. This process demands not only a precise understanding of fine-grained visual attributes, such as layout, structure, and spatial relationships, but also accurate alignment among images, textual semantics, and SVG code. Without explicit visual guidance, conventional cross-entropy supervision over code sequences struggles to capture complex visual semantics, often producing SVGs with limited visual fidelity or poor geometric consistency. To address these limitations, we propose **VA-SVG, a visual alignment post-training framework that couples SVG code prediction with visual generation.** Specifically, we train an auxiliary diffusion model to map continuous SVG code representations into the visual domain, and align its intermediate representations with visual features extracted from rendered SVGs. This differentiable pathway transfers visual generative knowledge back to the code generation model without requiring gradients through the SVG renderer. Extensive experiments and ablation studies demonstrate the effectiveness of the proposed training paradigm.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.