ChartVL: Benchmarking SVG Chart Generation through Visual Layer Reconstruction and Verifiable Judgment
Abstract
Evaluating SVG charts generated by large language models is challenging because appropriate layout, layering, and styling depend on the roles of chart elements and the task requirements. Existing approaches often use vision-language models (VLMs) to evaluate the visual quality of charts, but they are prone to missing or misjudging visual issues, and their judgments, expressed as scores or textual explanations, are difficult to verify. We introduce CHARTVL, a benchmark for SVG chart generation built on visual layer reconstruction and verifiable judgment. We reconstruct visual layers to inform VLM judgments and require executable conditions alongside each determinate visual-design judgment. Executing these conditions confirms or revises the initial verdicts using evidence from the rendered SVG. CHARTVL contains 360 tasks spanning four commonly used chart categories and three task levels with increasingly complex design requirements. Across four evaluated models, 83.1%–99.2% of outputs use the requested chart type and correctly present the required data and content, but only 9.7%–58.6% also pass the verified visual evaluation of layout, layering, and styling. Our findings show that chart generation requires more than correct content. Coordinating visual layers into a readable whole is a distinct challenge that calls for structured, verifiable evaluation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.