acceptodds
Under review as a conference paper at ICLR 2027

Vision2Code: A Multi-Domain Benchmark for Evaluating Image-to-Code Generation

Abstract

Image-to-code generation tests whether a vision-language model (VLM) can write executable code to reproduce an image. Existing benchmarks suffer in several ways, including narrow scope, dependence on source code, and generic evaluation criteria across very different domains. To address this, we introduce Vision2Code, a multi-domain and reference-code-free benchmark with 539 test examples from 15 datasets spanning six domains: charts and plots, geometry, graphs, scientific imagery, documents, and 3D spatial scenes. We also create a dataset-specific rubric that judges image reconstruction quality with emphasis on criteria relevant to that dataset. Taken together, our evaluation framework aligns better with human judgments from four annotators than generic VLM rubrics and pixel, embedding, and perceptual baselines. We study the robustness of our evaluation framework by computing agreement between two judges, evaluator repeatability, and the sensitivity to rubric weights and domain-specific errors. Results show that both VLM judges have low variance across repeated runs and preserve the same ranking of all nine evaluated models. Introduced perturbations are also correctly detected by both judges using our rubric. We show that benchmark performance varies substantially across domains. Models do well on charts and graphs but struggle with spatial scenes, chemistry, documents, and scientific diagrams. To improve image-to-code generation capability, we introduce evaluator-guided self-training, where evaluator scores are used to select model-generated programs for fine-tuning. This improves Qwen3.5-9B from 1.60 to 1.86 and Phi-4-MM from 0.58 to 0.74 on test, with gains also holding under additional evaluators. Vision2Code provides a multi-domain testbed for measuring, diagnosing, and improving executable visual reconstruction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.