GeoFlip: Do MLLMs Know Whether Geometry Problems Are Underdetermined?
Abstract
Advances in supervised fine-tuning and reinforcement learning have substantially improved the mathematical reasoning capabilities of multimodal large language models (MLLMs), enabling them to solve geometry problems when sufficient evidence is provided. However, whether MLLMs can reliably recognize underdetermination induced by removing essential visual evidence remains unclear. Existing reliability evaluations mainly focus on missing textual conditions, leaving the role of visual evidence in multimodal mathematical reasoning insufficiently explored. To address this gap, we evaluate MLLMs on paired solvable and underdetermined geometry problems, jointly assessing problem-solving ability and underdetermination recognition. Specifically, we propose Key-Information Removal (KIR), which uses source solutions to guide the removal of essential visual evidence from geometry images. To ensure the validity of the resulting problems, automatic problem auditing, targeted repair, and human verification are further applied so that the remaining conditions no longer uniquely determine the target. Through this process, we construct GeoFlip, a benchmark comprising hundreds of solvable–underdetermined geometry problem pairs. Experiments on GeoFlip reveal a substantial gap between problem-solving ability and underdetermination recognition, with models frequently repeating source answers even after necessary evidence has been removed. Moreover, underdetermination recognition varies substantially with the type and modality of removed evidence, while model scaling and reasoning modes do not consistently improve underdetermination recognition. Finally, to improve reasoning reliability in both solvable and underdetermined settings, we develop a unified symbolic reasoning pipeline that combines multimodal formalization with geometric deduction and algebraic computation. Compared with direct answering, our pipeline raises pair success rate from 0.3% to 59.5% with Claude Opus 5 and from 35.1% to 57.2% with GPT-5.5.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.