Learning an Independent Multimodal Auditor for Language-Grounded Segmentation
Abstract
Language-grounded segmentation models are increasingly used in multimodal systems, but they can produce confidently incorrect masks, including masks that select the wrong object entirely. This creates a distinct verification problem at deployment: given the image, referring expression, and a candidate mask, determine whether it covers the intended object accurately without seeing the ground-truth mask. A visually coherent segmentation can still be semantically wrong, so quality assessment must check the image–text–mask correspondence as well as mask geometry. Although much work improves mask generation, independent language-conditioned assessment of an already predicted mask remains underexplored. We introduce an independent multimodal large language model (MLLM)-based auditor that verifies predicted masks using only the image, referring expression, and mask itself, without access to ground truth or segmentation model-specific information. We formulate mask verification as estimating the quality of a predicted mask and show that language-conditioned attention provides a useful grounding signal to improve verification. The proposed auditor achieves a better balance between retaining correct masks and rejecting incorrect ones, outperforming existing mask-quality signals and learning-based baselines. It also generalizes across segmentation models and datasets, including challenging reasoning-based referring expressions. These results show that separating segmentation generation from verification provides an effective reliability layer for language-grounded vision systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.