UniSeek: Unified Referring Expression Localization with Count-Guided Instance Slots and Lightweight Add-ons
Abstract
Referring expression localization identifies a target set from natural language and represents its spatial extent using bounding boxes or masks. Unifying these tasks within a multimodal large language model (MLLM) requires handling variable target counts and heterogeneous geometric outputs while preserving fine spatial information. We propose UniSeek, a single model for referring expression comprehension (REC), referring expression segmentation (RES), and their generalized counterparts (GREC/GRES). UniSeek connects the MLLM backbone to downstream lightweight decoders through instance slots trained with explicit counting supervision. A Sparse Localization Decoder and a Dense Localization Decoder predict bounding boxes and masks, respectively. A Lightweight Image Local Detail Injector provides local detail features without an external pretrained image encoder. Experiments on both RefCOCO/+/g and gRefCOCO show that UniSeek achieves leading box-level localization results and competitive segmentation performance. For example, UniSeek-8B achieves a mean REC Acc@IoU=0.5 of 93.96% across eight splits, outperforming all compared baselines on every split. These evaluations clearly demonstrate that the proposed UniSeek re-calibrates the modern state-of-the-art performance under most settings of the interested tasks here.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.