acceptodds
Under review as a conference paper at ICLR 2027

Infographics as a Visual Language: Learning through Context and Reconstruction

Abstract

Advances in multimodal models have improved chart understanding and image generation, yet it remains unclear how well these models capture the relationships that make an infographic meaningful. Infographics communicate through a visual language in which numerical values, text, and graphical elements constrain and complement one another. We introduce **CoReL (Contextual Reconstruction Learning)**, a data construction and pretraining recipe for learning infographic visual language through contextual reconstruction. CoReL organizes learning around three complementary tasks—**Numerical Reasoning**, **Numerical Rendering**, and **Contextual Synthesis**—that connect numerical meaning, visual expression, and contextual compatibility. We build **InfoContext**, a training corpus of **1.7M images** with **27M numerical-text bounding boxes** and **22M graphical bounding boxes**. We also introduce **InfoContext-Bench**, a human-annotated benchmark spanning the three tasks. We apply CoReL to an existing unified understanding-and-generation model through element-masked reconstruction. Experiments reveal substantial limitations in the evaluated public models, even when additional information about the missing content is provided, while our model achieves stronger performance across all three tasks. Numerical-reasoning pretraining with CoReL further improves downstream chart question answering on both **BAGEL and Qwen**, with consistent gains across benchmarks. These results establish contextual reconstruction as an effective framework for learning infographic visual language, with benefits extending from element completion to reasoning across chart elements.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.