acceptodds
Under review as a conference paper at ICLR 2027

Found but Not Bound: Grounded Objects and Unfaithful Relations in Vision-Language Models

Abstract

We score each description of a compositional benchmark pair with one yes/no score and compare two criteria, whether the correct description ranks above the incorrect one, and whether each description gets the right answer on its own. On ARO, answering yes whenever "yes" is more probable than "no" reverses two model comparisons. A threshold fitted on held-out images removes both, yet 20.8 to 62.0% of correctly ranked pairs still contain a wrong standalone decision, and a smaller reversal appears. How much of the gap the threshold recovers varies by model, from most of it for InternVL3 on attribution to almost none for Qwen2.5-VL on relation. Across five models and six subsets, ranking exceeds standalone accuracy, and with both descriptions in one prompt the A/B scores of four models track ranking. Benchmarks that report only paired ranking can misorder models; we recommend also reporting standalone accuracy and the decision threshold behind it.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.