acceptodds
Under review as a conference paper at ICLR 2027

UniSeek: Unified Referring Expression Localization with Count-Guided Instance Slots and Lightweight Add-ons

Abstract

Referring expression localization identifies a target set from natural language and represents its spatial extent using bounding boxes or masks. Unifying these tasks within a multimodal large language model (MLLM) requires handling variable target counts and heterogeneous geometric outputs while preserving fine spatial information. We propose UniSeek, a single model for referring expression comprehension (REC), referring expression segmentation (RES), and their generalized counterparts (GREC/GRES). UniSeek connects the MLLM backbone to downstream lightweight decoders through instance slots trained with explicit counting supervision. A Sparse Localization Decoder and a Dense Localization Decoder predict bounding boxes and masks, respectively. A Lightweight Image Local Detail Injector provides local detail features without an external pretrained image encoder. Experiments on both RefCOCO/+/g and gRefCOCO show that UniSeek achieves leading box-level localization results and competitive segmentation performance. For example, UniSeek-8B achieves a mean REC Acc@IoU=0.5 of 93.96% across eight splits, outperforming all compared baselines on every split. These evaluations clearly demonstrate that the proposed UniSeek re-calibrates the modern state-of-the-art performance under most settings of the interested tasks here.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.