CaptureChart: How Reliable Is Chart Understanding across Capture Conditions?
Abstract
Reading structured visual information from photographed charts matters for automated field inspections, factory monitoring, and mobile question answering. As vision–language models advance, evaluation benchmarks must evolve to assess their growing capabilities and remaining limitations. Yet existing real-photo benchmarks are difficult to scale because adding new charts or capture conditions requires costly new photography. Natural captures also combine multiple degradations, making their individual effects on chart understanding difficult to isolate. Hence, we introduce CaptureChart, a fully digital benchmark for studying chart understanding under simulated capture. Starting from 1,000 CharXiv charts, a Blender-based pipeline models the paper surface, surrounding scene, lighting, and camera imaging to produce 32,000 aligned views. Each chart has an electronic original, a clean capture, and 30 variations across four capture families and 11 mechanisms at selected intensity levels. To examine difficult reasoning questions under capture degradation, we add 1,151 visual-evidence questions linked to 173 reasoning questions. Each probe targets a chart fact used in a checked solution to its parent question, allowing us to test whether models recover the evidence needed for the final answer. Across 14 vision–language models and 30 capture variations, mean accuracy falls relative to clean captures for descriptive, reasoning, and visual-evidence questions; reasoning accuracy drops by 9.4 percentage points. Reasoning is often more sensitive than descriptive reading. Linked evidence answers can change correctness even when aggregate accuracy changes little, and specific capture mechanisms expose distinct model weaknesses. Overall, CaptureChart offers a scalable way to study chart understanding under controlled, simulated capture conditions. The benchmark is available at https://anonymous-hf.com/a/7if3j4z0hjpy/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.