CompVis-Bench: Benchmarking Structural Understanding of Composite Visualizations
Abstract
Many real-world visualizations combine charts, maps, and glyphs whose compo- nents are connected through shared data fields. Understanding these composite visualizations requires recovering not only each component’s visual encoding, but also the data relationships that connect the components. We introduce COMPVIS- BENCH, a reconstruction-based benchmark for evaluating this structural under standing in vision-language models. Its normalized mark-group representation preserves data semantics, encoding structure, and cross-group field correspon dences while abstracting away presentation details and alternative composition hierarchies. We build an expert-annotated seed set and extend it through con trolled structural modifications, randomized rendering, and manual verification, yielding 1,600 verified visualizations. Across 12 large vision-language models, local mark groups and encodings are recovered substantially more reliably than the field correspondences that connect them. These results expose a central lim itation of current models: recognizing the parts of a complex visualization does not imply understanding how those parts jointly represent data.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.