acceptodds
Under review as a conference paper at ICLR 2027

StructViz-Bench: A Controlled Study of Visualization-Format Dependence in Multimodal LLM Reasoning over Structured Data

Abstract

Multimodal LLMs increasingly reason over structured data that reaches them as a picture—a bar chart, a heatmap, a node–link diagram—and the rendering is usually chosen incidentally by whatever tool drew the figure. We ask whether that arbitrary choice changes what the model concludes, and introduce StructViz-Bench, which treats visualization format as a controlled variable—fixing the underlying data and the question semantics and varying only how the data is drawn—and is, to our knowledge, the first to do so across structured modalities: tabular, time-series, and graph under one protocol. The benchmark comprises 3,795 base items rendered into 14 visualization types over four reasoning-depth categories. Evaluating several proprietary and open MLLMs, we establish that the choice matters, and by a wide margin: within-modality best–worst exact-match gaps reach 40.5pp, roughly half of all base questions (45–56% per model) change correctness under some format swap, and paired item-level tests are significant for all 12 model×modality comparisons (Bonferroni-corrected p<0.01). The effect is not an artifact of the scoring rule or of layout randomness, and it persists across model families, a 4.6× parameter increase, and all four prompt styles we test—including chain-of-thought, which never helps (4.4–14.6pp below each model's best non-CoT prompt). An answerability audit of every (question, rendering) pair then shows what these gaps are made of: about half of all pairs do not encode the answer at all (a bar chart of column means has no rows; a drawing of a large graph has no node labels), and on those, accuracy sits at prior level for every model. Re-rendering the suite so that every format encodes the answer and re-evaluating an open model, the answer-less renderings recover 16–23pp, the graph listing loses 10pp once it stops printing node degrees, and the rendering effect that remains for the same information is 11pp on tables and 5pp on graphs; none is identified for time series, where no rendering but the complete listing beats a question-only prior, and no graph rendering beats it at all for this model. Printing the queried quantity matters more than the chart type. For practice, much of the accuracy cost is recoverable by choosing the rendering: a learned lookup that picks the best rendering per question type, with no weight update, recovers up to 19.5pp over a random format and 2–8pp over the best fixed format per modality (object-level CIs exclude zero for all seven models; a hand rule that only picks an encoding format captures 0.4–2.7pp of it); and multi-format fine-tuning raises cross-format agreement in every modality (Consistency Rate up 15–25pp) and narrows the accuracy gap in two of three, while a λ=0 control shows that a consistency regulariser adds agreement but no accuracy, partly by making errors consistent. We release the benchmark, the deterministic generation pipeline, and all evaluation and mitigation code.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.