GeoPerturb: Probing Structural Vulnerability in Multimodal Geometry Reasoning
Abstract
Large Multimodal Models achieve strong results on visual geometry benchmarks, yet it remains unclear whether this performance reflects geometric reasoning that remains stable across equivalent drawings. Rather than relying on low-level pixel corruptions, we introduce GeoPerturb, a matched diagnostic framework that reconstructs each problem in TikZ and generates four answer-preserving variants that change orientation, symbol identity, visual appearance, or irrelevant structure. The resulting benchmark contains 531 matched problem suites. Across proprietary and open-source LMMs, we find that the same problem is often answered correctly in one equivalent view and incorrectly in another. This instability is selective rather than generic: rotation and relabeling expose the clearest failures, while appearance changes and irrelevant clutter are less disruptive for frontier models. Marginal accuracy can further conceal these failures because gains on previously incorrect problems can offset losses on problems already solved; canonical-conditioned retention makes these solved-to-failed transitions visible. A controlled fine-tuning study further shows that explicit derivation supervision, rather than exposure to additional views alone, drives cross-view improvement. GeoPerturb therefore shifts geometry evaluation from asking whether a model can answer one drawing to asking whether it preserves the same solution across equivalent drawings. The central weakness is not simply low accuracy, but unstable correctness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.