acceptodds
Under review as a conference paper at ICLR 2027

RefCon: Learning What to Include & Exclude for MLLM Segmentation

Abstract

Multimodal large language models (MLLMs) have enabled language-guided image segmentation, including referring expression segmentation. However, MLLM-based segmentation methods remain referent-centric; when exclusion information is used, it is typically encoded as sparse point prompts rather than explicit non-referent image-patch representations maintained throughout decoding. Consequently, explicit non-referent image-patch evidence generally remains absent from the MLLM-generated representation used for dense prediction, leaving the referent represented primarily through what should be included. We introduce RefCon, a framework that represents each referent through complementary MLLM-generated referent and non-referent image-patch representations and maintains both throughout decoding for repeated contrast. Given an image and instruction, the MLLM autoregressively generates two image-grounded token sets: Referent Tokens, corresponding to referred-object patches, and Contrast Tokens, corresponding to non-referent patches. To exploit this complementary evidence, we introduce a Referent–Contrast Decoder (RCD) that repeatedly refines both branches with image features and explicitly contrasts their token and dense-feature representations throughout refinement. To structure these token sets for role separation and non-redundant coverage, we further propose a Region-Aware Contrastive Transport (RCT) loss: its contrastive component aligns Referent Tokens with foreground evidence and Contrast Tokens with non-referent evidence while separating their roles, whereas its optimal-transport component reduces within-set redundancy and encourages collective coverage without assigning predefined subregions to individual tokens. Together, RefCon introduces explicit referent–non-referent representation into MLLM segmentation, with complementary inclusion and exclusion evidence generated by the MLLM and repeatedly contrasted throughout dense decoding.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.