acceptodds
Under review as a conference paper at ICLR 2027

DeepVisualMath: Comprehensive Evaluation of Visual Mathematical Reasoning in Large Multimodal Models

Abstract

Visual mathematical reasoning requires integrating visual evidence with mathematical concepts, yet its evaluation in multimodal large language models (MLLMs) remains fragmented across benchmarks targeting limited domains and skills. We introduce ***DeepVisualMath***, a unified benchmark and evaluation framework for holistic assessment of visual mathematical reasoning. The benchmark comprises **10,298 questions**, each paired with one image, organized into **7 categories, 18 subcategories, and 80 tasks**, spanning applied sciences, continuous mathematics, relational geometry across dimensions, discrete mathematics and formal logic, plane geometry, contextual mathematics, and solid geometry. Beyond the benchmark, we provide an automated curation pipeline, a reproducible evaluation harness, and an automated error analysis framework for detailed diagnosis across the task hierarchy. Evaluation of 10 MLLMs reveals substantial and uneven performance gaps: overall accuracy ranges from 15.50% to 66.67%, with even the strongest model leaving approximately one third of questions unresolved. Performance also varies markedly across mathematical domains; the leading model achieves 77.36% in continuous mathematics but only 51.85% in plane geometry, highlighting limitations obscured by aggregate scores. These findings demonstrate that strong performance in one mathematical domain does not imply broad visual mathematical competence. ***DeepVisualMath*** provides a systematic foundation for evaluating these capabilities and identifying priorities for improving mathematical reasoning grounded in visual evidence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.