CROSSREF-VQA: Benchmarking Cross-modal Reference Resolution in Structured Graphics
Abstract
While MLLMs perform well on everyday visual question answering (VQA) tasks, they remain unreliable when required to combine visual and textual information at the level of individual elements. Interpreting structured graphics demands exactly this: elements are marked with short references such as legend keys and colour codes with the corresponding semantic identities in a separate text section, leaving the reader to cross-reference. To study this limitation, we propose CrossRef-VQA, a benchmark explicitly designed to isolate element-precise cross-modal reference resolution. CrossRef-VQA renders structured data as graphics, replaces every semantic label in the image with an anonymous reference, supplies the reference-to-element mapping as text, and computes every answer programmatically from the underlying data. We intensionally incorporate human designed templates to avoid circular LLM evaluation and support scalable data generation, enabling stable large-scale assessment. Unlike traditional VQA benchmarks that evaluate a single, confounded setup – where full semantic information is redundantly present in both the image and the text – our framework introduces a set of controlled experimental conditions. By keeping the anonymised graphic fixed and systematically varying only the need to perform cross-modality referencing and de-referencing, we can directly compare performance across conditions to cleanly isolate cross-modal integration from pure visual perception. Across 8 chart types and 11 MLLMs, we find that requiring cross-referencing causes a significant overall performance drop of 15–30%. Notably, the two directions fail differently: referencing is imprecise when the question specifies an element, whereas de-referencing is skipped, where models correctly locate the visual element but return its raw label most of the time. These findings demonstrate that the fundamental failure point is not visual understanding, but the cross-modal reference resolution itself.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.