Single-Domain Generalizable Open World Object Detection
Abstract
Open world object detection (OWOD) extends closed-set detection by identifying unknown objects and incrementally expanding the known class set. However, its assumption that training and evaluation data share the same domain limits applicability in real-world deployment, where visual domains continually change. Our experiments showed that even a state-of-the-art OWOD detector suffered performance degradation under such shifts, and a straightforward combination with a domain-generalized object detection (DGOD) method provided only marginal recovery. To bridge this gap, we introduce **Single-Domain Generalizable Open World Object Detection (SG-OWOD)**, where a detector trained on a single source domain must preserve its open world capability across multiple unseen target domains. To investigate these challenges, we analyzed a vision-language-based OWOD baseline. We found that domain shifts were associated with feature drift across the backbone and weaker known-class responses. Its attribute-selection statistics were also dominated by background, obscuring object-related evidence for unknown discovery. These observations motivate **RADAR** (**R**obust **A**ttribute-guided **D**etection via **A**ligned **R**epresentations), a framework built on a pretrained vision-language object detector (VLOD) with two complementary components. *Consistency-driven Low-rank Adaptation (CLA)* adapts early backbone layers by enforcing prediction consistency under style perturbations, improving known-class detection under domain shifts. *Background-Decoupled Attribute Selection (BDAS)* builds an object-centric attribute bank by separating object-like regions from background during selection, improving unknown recall when combined with CLA. Across three target domain benchmarks, RADAR improved over the strongest published baselines by at least K-mAP at Task 4 and H-Score at Task 3, while outperforming a naive OWOD and DGOD combination at every incremental task.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.