WHAT DOES SPATIAL ACCURACY CERTIFY? MATCHED AND CROSSED CANDIDATE AUDITS
Abstract
Spatial multiple-choice accuracy can conceal candidate shortcuts and unstable an- swers under controlled layout changes. We test what such scores certify using question-blind pixel rules, crossed scene–instruction cases, matched foil transla- tions and repeated responses across six model configurations and 47,136 formal requests. A fixed question-blind rule solves 245/326 released STARE questions, though candidate-only prompting performs poorly on a matched subset. Our bal- anced construction gives exact 1/4 accuracy ceilings for specified input-omission rules. Yet Qwen2.5-VL-32B (Qwen32) original preprocessed inputs admit a per- fect query-plus-centroid rule; direct-grid rendering restores seven declared source and native bounds. On one aligned 78-cell shape, Qwen32 rejects both omission nulls, but 54.27% of cross-layout response pairs disagree, versus 2.05% of same- input pairs. An exact categorical permutation test rejects response-distribution equality; the mean half-squared categorical distance is 0.5222 (95% interval [0.4586, 0.5859]). Separately, the simultaneous 95% interval for Qwen32’s disk- pair mean-accuracy change is [−3.18, 3.48] points, within the prespecified ±5- point margin despite 397 correctness reversals among 2,048 pairs. These results separate evidence for required inputs, mean tolerance and response stability with- out identifying an internal spatial algorithm.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.