HARE: Hierarchy-Aware Routing of LoRA Experts for Open-Vocabulary Object Detection at Higher Semantic Levels
Abstract
Open-vocabulary object detectors perform strongly on concrete categories but often miss relevant instances when queried with higher-level concepts such as animal, vehicle, or furniture. We study higher-level category detection, which requires localizing instances from visually diverse concrete subclasses using their shared ancestor concept as the query. To support training and evaluation, we construct hierarchy-annotated versions of COCO, LVIS-mini, and Objects365 by associating each source category with a concrete-to-abstract semantic path and propagating its instance annotations to the concepts on that path without additional bounding-box annotation. We further propose Hierarchy-Aware Routing of Experts (HARE), a simple parameter-efficient adaptation framework that uses parent-child relations as explicit supervision for allocating adaptation parameters across category queries. Specifically, a ranking objective encourages parent concepts to receive higher abstraction scores than their children, guiding a text-conditioned router over lightweight LoRA experts. HARE is adapted only on Objects365 and evaluated on held-out Objects365 categories and the unseen COCO and LVIS-mini datasets. Under direct higher-level queries, HARE improves average box AP over Grounding DINO by 8.5, 5.1, and 4.2 points on COCO, LVIS-mini, and Objects365, respectively, while matching its concrete-category AP on COCO. Under identical subclass augmentation, HARE retains gains of 3.7, 2.3, and 1.1 points. These results demonstrate that explicit hierarchy supervision improves higher-level detection and transfers across datasets, with substantial gains even without auxiliary subclass cues at inference.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.