Where Grounding Accuracy Lives on the IoU Curve: Label-Free Inference-Time Boundary Refinement
Abstract
Vision–language models can identify the correct referent while returning an imprecise bounding box. We study whether a frozen direct-answer model can use its own prediction to allocate and safely exploit a localized re-observation without accessing target annotations at inference. Label-free precision refinement (LFPR) routes predicted-small regions to a higher-resolution pass, re-grounds the expression inside a context crop, admits a candidate only under fixed geometric guards, and returns a fixed coordinate-wise midpoint. We report results across three evidence tiers. On 31,921 Ref-L4 expressions – retrospective evidence, since the published test distribution informed method development – LFPR raises our primary endpoint, mAcc(0.5:0.95), from 72.947% to 76.013% ([email protected] 88.531% to 89.725%, [email protected] 55.788% to 61.142%). A frozen, no-retuning transfer to 30,969 RefCOCO/RefCOCO+/RefCOCOg expressions – a later content-hash audit found substantial image overlap with Ref-L4 – improves every dataset individually at [email protected], mAcc, and mean IoU (mAcc +0.645, [email protected] +0.817, mean IoU +0.552 points pooled), while [email protected] is not distinguishable from zero on any dataset or pooled – a stage-wise audit shows routing alone gains +1.162 points at [email protected] while crop, guards, and fusion give back -1.192 points relative to that routed arm, so the two stages substantially offset rather than the operator having no strict-IoU effect. A prospective, image-disjoint Flickr30K Entities evaluation improves every endpoint under an MDETR-style merged-box protocol (mAcc +0.973, [email protected] +1.022) and more strongly under a historical single-box selection on the same images (mAcc +2.575, [email protected] +3.689). Applying the identical operator to two released grounding specialists improves every reported endpoint ([email protected] +1.569 for EGM-4B, +6.716 for EGM-8B) at roughly twice the inference latency, showing that the refinement composes with specialist training rather than replacing it. On the untuned transfer and prospective tiers, a genuine unguarded control – the same candidates with the guard removed – underperforms the incumbent, showing the guard is load-bearing. Together, these results show that referent selection and boundary precision are partially separable, and that different inference components move different, sometimes opposing, regions of the IoU curve – behavior a single accuracy threshold cannot reveal.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.