Counterfactual Twins and Executable Validity Gates for Diagnosing Spatial Reasoning in Multimodal LLMs
Abstract
Benchmark accuracy reports how often a multimodal LLM answers correctly, not whether it answered by looking. We present TWIN-Diag, a spatial-reasoning benchmark built from counterfactual twins. Each item pairs a procedurally generated indoor scene with two single-object interventions of identical displacement magnitude, one flipping the ground-truth answer and the other preserving it. Across eleven MLLMs from five families, answers track the image but not the truth. On relative distance (T-5), the one template passing every gate, answers change on 21.6% of controls, where the truth is held, and 24.1% of treatments, where it flips, against 5.0% for a repeated image. Pooled over models, counterfactual sensitivity is at most +0.069 (one-sided 95% bound, rooms as the unit of inference), and no model-template cell survives correction. The control's construction matters. In a pre-registered test on real renders of our previous release's pairs, controls sampled to keep the answer changed fewer answers than controls solved to the treatment's exact magnitude in all ten models (room-level p = 0.019, Holm-rejected), inflating measured sensitivity. Ours also match on-screen travel within 10%. Nine executable validity gates read only the data and exit non-zero on failure. All were fixed before any model answered the release (78 pairs, 32 rooms), four after models had answered earlier builds. The gates are binding. An extended gate rejected our previous release, and absolute distance (T-2) fails the question-count gate, so it is descriptive. The plan's confirmatory test, an image ablation on six unseen T-2 items, rejects at its design floor (p = 1/81), in conflict with that gate. On T-5 the same ablation fails to show that the image improves accuracy. Physical-plausibility defects that no gate checks do not detectably separate the arms. We group the twenty-four defects found in our work into seven failure modes of counterfactual-benchmark construction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.