UniDiagCode: A Multi-Format Benchmark for Image-to-Code Diagram Generation
Abstract
Image-to-code diagram generation converts rasterized diagrams into editable code representations. Prior work largely targets a single format, despite substantial differences in expressivity, code complexity, and downstream compatibility. We introduce UniDiagCode, a unified six-format benchmark comprising 341K real-world rendered-image/code pairs and 6.3K hand-drawn sketches. We evaluate large vision language models across formats and reveal substantial differences in visual fidelity, renderability, and generation cost. We further formulate target-format selection as an input-dependent decision, allowing the model to choose an appropriate format for each diagram. This model-driven selection achieves a better quality–cost trade-off than fixed-format generation, while fine-tuning on our data substantially improves performance under both format specification and selection. By treating the output format as a model decision rather than a predefined constraint, our work offers a flexible and practical framework for code-based diagram generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.