acceptodds
Under review as a conference paper at ICLR 2027

Towards Less Base-Biased Open-Vocabulary Object Detection

Abstract

Frozen-backbone detectors have emerged as an effective paradigm for Open-Vocabulary Object Detection (OVOD), as they preserve transferable knowledge from pretrained vision-language models. However, freezing the backbone alone does not eliminate detector-level base bias: base-only optimization can distort novel-region representations, while region proposal networks trained on base categories often produce spatially inaccurate proposals for unseen objects. To address these issues, we propose ToLB, a framework Towards Less base-Biased OVOD that tackles representation and localization bias at training and inference, respectively. Specifically, we propose a Relational Knowledge Preservation module that employs a frozen visual foundation model as a category-agnostic structural teacher and regularizes region representations through pairwise relational consistency, avoiding restrictive alignment between heterogeneous feature spaces. At inference, we further introduce a VLM-Guided Region Refinement module that extracts category-agnostic objectness cues from dense vision-language features to identify foreground-dominant regions and adaptively refine proposals, thereby reducing background interference for novel objects. Extensive experiments on OV-COCO and OV-LVIS demonstrate superior performance in novel-category generalization, achieving 48.2% Novel AP50 on OV-COCO and 39.5% Rare mask mAP on OV-LVIS, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.