OpusDet: A Simple yet Effective Unified Framework for Open-Vocabulary Detection
Abstract
Recent unified open-vocabulary detection (OVD) methods support text queries, visual exemplars, and their combinations. However, they often rely on increasingly complex designs such as heavy cross-modal fusion, staged training, and iterative annotation pipelines. We revisit whether such complexity is necessary with stronger foundation models available. We find that unified OVD can be made much simpler with strong visual features and scalable grounding supervision. We present OpusDet (Open-vocabulary, Prompt-Unified, Simple), a unified detector supporting text, interactive visual, generic visual, and mixed prompting. OpusDet keeps the model, training, and data pipeline simple. It uses one detector for all prompt types, one-stage text–visual training, and a SAM3-based single-pass data engine. From the matched Base to the Full data setting, with the same Swin-T backbone, OpusDet gains 6.5 Visual-I AP on LVIS-minival, compared with 3.1 AP for our T-Rex2-style re-implementation. Our final DINOv3-ConvNeXt-B model achieves state-of-the-art Visual-I performance among comparable unified detectors, reaching 68.1, 69.2, and 54.7 AP on COCO, LVIS-minival, and ODinW35. It also maintains competitive Text and Visual-G performance. It further turns mixed prompting from interference into complementarity, improving over text or visual prompts alone. These results show that a simple unified detector can provide strong performance across prompting modes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.