acceptodds
Under review as a conference paper at ICLR 2027

GroundBench: Multi-Resolution Polygon Grounding Exposes the Geometry Gap in Vision-Language Models

Abstract

A bounding box can locate the right object without describing its shape. This leaves an important question unanswered by box-based grounding scores: can a general-purpose vision–language model turn the same visual and linguistic evidence into a usable object region? We introduce GroundBench, a benchmark that asks models to express referred objects as polygons through their text-generation interface. It pairs 1,500 image–expression–referent triples across five vertex budgets, holding the questions fixed while varying output complexity. Evaluation reports filled-region overlap, coordinate-sequence completion, and polygon legality separately. Across four hosted models, more vertices do not reliably improve the output: performance degrades at the largest budget despite more faithful reference contours. Target preference can also remain high while conditional region quality is poor, showing that favouring the right object does not ensure a useful geometric description. Controlled interventions further reveal sensitivity to reasoning settings and misleading spatial language. These findings expose a geometric-output gap: locating an object does not ensure reliable explicit geometry. They identify improvement targets in region construction and output completion that a box score alone does not expose. GroundBench evaluates their combined outcome rather than latent boundary perception in isolation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.