acceptodds
Under review as a conference paper at ICLR 2027

KNOWING HOW MANY IS NOT KNOWING WHICH ONES:NATIVE COMPLETE-SET GROUNDING IN MULTIMODAL LARGE LANGUAGE MODELS

Abstract

Generalized Referring Expression Comprehension (GREC) requires a model to recover all objects referred to by an expression, including no-target and multi-target cases. We identify a spurious-correlation failure mode in GREC: target cardinality remains partly predictable when visual evidence is removed or mismatched, revealing reliance on language and dataset priors. Yet accurate counting does not imply correct target selection. A strong query-based set decoder predicts the correct count on 98.38% of two-target expressions but achieves only 10.90% Pr for the complete target set. When the query changes to refer to a different pair of objects in the same image, the decoder usually preserves the count but rarely switches to the correct targets. We therefore propose Native Complete-Set Grounding, which trains the MLLM to select and localize all referred objects through its autoregressive localization pathway. In a matched comparison with a separate set-prediction decoder, Pr improves from 50.24% to 77.22%. Our 3B model achieves 81.24/76.18/68.11% Pr on gRefCOCO val/testA/testB, establishing state-of-the-art performance on all three splits. The gains persist at 7B scale while conventional REC performance remains strong.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.