acceptodds
Under review as a conference paper at ICLR 2027

GEOMETRIC PARROTS: RULE ABSENCE VS. RULE FRAGILITY IN SPATIAL AND PHYSICAL REASONING FOR VISION- LANGUAGE MODELS

Abstract

Large language models have been characterised as stochastic parrots- fluent producers of plausible text without grounded understanding of what that text describes. We ask whether the same holds for vision-language models, and observe that accuracy cannot answer the question. A model that answers a geometric question correctly and then abandons that answer the moment a user asserts otherwise did not hold the rule; a model that answers wrongly and revises toward the truth is reasoning well. Accuracy scores these identically, because it evaluates each response in isolation. What separates them is the direction in which an answer moves under challenge, and direction is egible only where the correct answer is known exactly — which annotated benchmarks, whose labels are human judgements about pre-existing images, cannot provide. We present GRIP, a suite of 34 procedurally generated domains comprising 100,000 images and 600,000 questions spanning plane, solid, projective, transformational, topological, analytic, inductive, optical and physical reasoning. Every scene is drawn by deterministic code and every answer computed in closed form from the coordinates the renderer itself uses; ground truth is derived rather than annotated, and no generative model participates at any stage. On this foundation we introduce Dual-Loop Evaluation. The closed loop poses five ordered questions per image on a compositional ladder in which each level consumes the previous level's result, localising where competence stops. The open loop then challenges a free-response answer twice over: under cross-examination, where the model is shown the answers other models gave, and under answer sycophancy, where a user asserts an correct answer without evidence. Because ground truth is exact, every resulting change of answer is signed as a correction or a capitulation — a decomposition that existing work on answer instability reports only as an unsigned flip rate. Rule absence and rule fragility thus become separately measurable: the first as failure in the closed loop, the second as a correct answer that does not survive challenge.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.