acceptodds
Under review as a conference paper at ICLR 2027

UniOSE: Semantics-Enriched Object Queries for Open-Ended Detection and Segmentation

Abstract

Open-ended object detection and segmentation require discovering objects and generating their category names without a predefined category vocabulary. Existing methods typically generate names from object representations produced by a detector. We find that insufficient category semantics in these representations limit naming accuracy, while enriched object semantics can further improve detection and segmentation. We therefore propose UniOSE, a unified framework that constructs semantics-enriched object queries to support both category generation and spatial prediction. Our Semantics-Enriched Object Namer combines detector features with region-level semantics from a pretrained vision–language model, providing richer semantic representations for category generation. Building on these representations, the Semantics-Guided Object Detector incorporates semantic compatibility between object queries and target categories into training-time matching, allowing supervision assignment to account for both object semantics and spatial location. The Semantically Conditioned Mask Decoder injects the same queries into mask decoding, enabling object semantics and local visual evidence to jointly inform mask prediction. Experiments demonstrate strong performance in open-ended object detection and segmentation, including a 5.0 AP gain over the prior method on LVIS minival.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.