acceptodds
Under review as a conference paper at ICLR 2027

Learning an Independent Multimodal Auditor for Language-Grounded Segmentation

Abstract

Language-grounded segmentation models are increasingly used in multimodal systems, but they can produce confidently incorrect masks, including masks that select the wrong object entirely. This creates a distinct verification problem at deployment: given the image, referring expression, and a candidate mask, determine whether it covers the intended object accurately without seeing the ground-truth mask. A visually coherent segmentation can still be semantically wrong, so quality assessment must check the image–text–mask correspondence as well as mask geometry. Although much work improves mask generation, independent language-conditioned assessment of an already predicted mask remains underexplored. We introduce an independent multimodal large language model (MLLM)-based auditor that verifies predicted masks using only the image, referring expression, and mask itself, without access to ground truth or segmentation model-specific information. We formulate mask verification as estimating the quality of a predicted mask and show that language-conditioned attention provides a useful grounding signal to improve verification. The proposed auditor achieves a better balance between retaining correct masks and rejecting incorrect ones, outperforming existing mask-quality signals and learning-based baselines. It also generalizes across segmentation models and datasets, including challenging reasoning-based referring expressions. These results show that separating segmentation generation from verification provides an effective reliability layer for language-grounded vision systems.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.