Categories as Prototype Sets: Constructing Text Prototypes from Descriptions for Open-Vocabulary Object Detection
Abstract
Open-vocabulary object detection uses text representations to recognize categories that have no bounding-box annotations during training. Existing approaches often represent each category with a class-name embedding or a fixed average of description embeddings, limiting how descriptions contribute to region-level recognition. We propose ProtoSet-DETR, a DETR-based detector that learns multiple text prototypes per category from a fixed set of offline-generated descriptions. A shared Text Prototype Aggregator constructs the prototypes from frozen CLIP description embeddings using attention and a learned value projection. The detector uses this bank for proposal scoring, decoder query initialization, and classification, weighting prototype contributions according to each visual region. Adaptive Prototype Regularization encourages diversity and balanced usage, while a training-only region–prototype alignment objective connects the bank to clustered encoder features. The same aggregator constructs prototypes for novel categories. Our diagnostics show concentrated description attention and reduced redundancy among the projected embeddings. On LVIS, ProtoSet-DETR achieves 46.0 rare-category AP (APr) and 46.1 AP, improving over LaMI-DETR with the same backbone by 2.6 APr and 4.8 AP. Without fine-tuning, the LVIS-trained model transfers to COCO and Objects365 with 46.3 and 25.4 AP. These results support learning how category descriptions are represented and matched to regions. Code will be released here.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.