Benchmarking and Diagnosing Continuous Geometric Reasoning in Vision-Language Models
Abstract
Vision-language models (VLMs) perform well on many spatial benchmarks, but their ability to reason over continuous geometry remains unclear. We introduce two complementary tasks: TangramPose, requiring object-specific prediction of position, rotation, and scale, and CanvasPlace, requiring placements that satisfy geometric constraints. Across 10 VLMs, most models achieve reasonable scene overlap yet fail strict object-level reconstruction, while the strongest model per- forms substantially better in small number piece task but fail when piece number increase. In-context demonstrations mainly improve scale calibration, whereas fine-tuning yields much larger gains, even with a frozen visual encoder. However, these gains transfer poorly to disjoint rotation ranges, unseen piece counts, and a distinct placement task. Probing shows that angle and scale are decodable from the base models’ visual representations far more accurately than they are expressed in final predictions. Counterfactual and activation-patching analyses further show that adaptation improves object-specific responses and strengthens the influence of target-region states on geometric outputs. Together, our results distinguish accurate in-distribution geometric prediction from robust, reusable geometric competence, and provide a framework for analyzing where VLMs succeed and fail in continuous geometric reasoning. Dataset and code will be released upon acceptance
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.