acceptodds
Under review as a conference paper at ICLR 2027

PictureProof: Evaluating Mathematical Reasoning in Vision-Language Models with Visual Proofs

Abstract

Human mathematical reasoning is profoundly visual at the level of intuition, discovery, and communication. However, existing visual math benchmarks evaluate on shallow calculation-based tasks rather than deeper modes of reasoning. To address this benchmarking gap, we introduce , a benchmark of visual mathematical reasoning based on translating visual math proofs into natural language. Because the grading for this task is qualitative, we implement a framework in which VLM outputs are graded at scale with an LLM-judge using human expert-crafted rubrics for each benchmark item. Our evaluations reveal dramatic failures in visual mathematical reasoning both on open-source as well as frontier VLMs. We further demonstrate that some of the failures are caused by a decay in attention to visual tokens during reasoning. We open-source our benchmark — specifically, the visual math proofs in TikZ code format, rubrics, and LLM-judge prompts — in order for the community to run further experiments on mathematical reasoning in VLMs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.