acceptodds
Under review as a conference paper at ICLR 2027

WHAT DOES SPATIAL ACCURACY CERTIFY? MATCHED AND CROSSED CANDIDATE AUDITS

Abstract

Spatial multiple-choice accuracy can conceal candidate shortcuts and unstable an- swers under controlled layout changes. We test what such scores certify using question-blind pixel rules, crossed scene–instruction cases, matched foil transla- tions and repeated responses across six model configurations and 47,136 formal requests. A fixed question-blind rule solves 245/326 released STARE questions, though candidate-only prompting performs poorly on a matched subset. Our bal- anced construction gives exact 1/4 accuracy ceilings for specified input-omission rules. Yet Qwen2.5-VL-32B (Qwen32) original preprocessed inputs admit a per- fect query-plus-centroid rule; direct-grid rendering restores seven declared source and native bounds. On one aligned 78-cell shape, Qwen32 rejects both omission nulls, but 54.27% of cross-layout response pairs disagree, versus 2.05% of same- input pairs. An exact categorical permutation test rejects response-distribution equality; the mean half-squared categorical distance is 0.5222 (95% interval [0.4586, 0.5859]). Separately, the simultaneous 95% interval for Qwen32’s disk- pair mean-accuracy change is [−3.18, 3.48] points, within the prespecified ±5- point margin despite 397 correctness reversals among 2,048 pairs. These results separate evidence for required inputs, mean tolerance and response stability with- out identifying an internal spatial algorithm.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.