MathVis: Can AI Reason through Drawing? A Benchmark for Mathematical Figure Generation
Abstract
This paper introduces MathVis, a benchmark designed to evaluate LLMs' ability to visualize frontier mathematical problems and perform diagram-assisted reasoning. The benchmark consists of 512 tasks, each comprising a mathematical statement, a reference diagram, executable TikZ/SVG code, a Gold Intermediate Representation (IR), atomic rubrics, and error codes. Through rigorous evaluation, GPT-6-astra, Claude Fable 5, and Gemini 3.1 Pro achieve Pass@1 scores of merely31.25%,26.56%,and22.85%,respectively. We found that although the tasks require visualization, LLMs still reason primarily through formal languages and symbols, rarely relying on visual diagrams. Furthermore, while the solutions generated by LLMs may contain all the necessary elements, they are often visually confusing to humans. Ultimately, this benchmark aims to visualize mathematical problems and promote research on enabling LLMs to simplify their reasoning via visual diagrams.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.