TFAG-OVD: Training-Free Attribute-Guided Open-Vocabulary Detection via LLM Semantic Composition
Abstract
Currently, open-vocabulary detectors recognize novel concepts primarily by aligning region features with category-name embeddings. However, as the vocabulary expands to thousands of categories, semantically related names crowd the alignment space, making the top match uncertain even at high similarity and disproportionately harming rare or weakly grounded concepts. Therefore, we rethink recognition as attribute-guided object understanding, representing ambiguous regions through transferable, class-agnostic attributes whose compositions can describe objects across datasets and vocabularies. A framework termed Training-Free Attribute-Guided Open-Vocabulary Detection via LLM Semantic Composition (TFAG-OVD) is proposed around a reusable pool of 3,549 class-agnostic attributes. Specifically, a frozen detector produces grounded proposals, and distribution-aware verification preserves clear predictions while routing ambiguous ones for attribute-guided review. For each routed region, a frozen vision-language recognizer retrieves shape, material, part, surface, and state cues; a text-only large language model then composes them within a task-relevant category range to revise the label. This selective attribute interface converts rigid region-to-name matching into evidence-based reasoning while keeping all models frozen. On LVIS minival, TFAG-OVD improves AP from 41.4 to 43.8 and rare-category AP from 34.1 to 37.8 over the frozen detector baseline. Zero-shot results on LVIS val, referring-expression benchmarks, ODinW, COCO, and COCO-O further demonstrate transfer across taxonomies and domains without parameter updates.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.