CounterGround: Diagnosing and Repairing Counterfactual Failures in Visual Grounding
Abstract
Visual grounding is typically evaluated by box overlap, which conflates association with the intended instance and localization precision. We introduce CounterGround, a paired counterfactual benchmark that fixes the image, object category, and candidate set while a spatial relation edit switches the annotated target instance. Across 5,635 pairs from remote-sensing and natural-image domains, recent multimodal large language models (MLLMs) exhibit a consistent behavioral separation: predictions can respond consistently to the relation edit and remain associated with the intended instances while failing the paired intersection-over-union (IoU) criterion. This separation persists as Qwen3.5 scales from 2B to 9B, despite higher grounding accuracy. Controlled Diagnosis on Qwen3.5-2B and Qwen3-VL-2B-Instruct shows that selecting among explicit ground-truth candidates is substantially easier than free-form box localization, while fixed-center extent prediction and a ground-truth-size oracle leave substantial residual error. Guided by this diagnosis, joint geometric refinement raises Pair Accuracy on CounterGround-DIOR validation from 48.65% to 56.06%, and adding local visual-language features raises it to 61.33%. The repair formulation also improves the reported point estimates on held-out CounterGround-DIOR with both backbones and on CounterGround-COCO, RefCOCO, and VRSBench under their specified training protocols. CounterGround links controlled target-switch evaluation to behavioral diagnosis and targeted repair.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.