MGME: Manifold-Guided Multi-Prototype Expansion for Open World Object Detection
Abstract
In Open World Object Detection, a detector must recognize known classes while simultaneously identifying instances of unknown categories. Existing approaches typically represent unknown objects using the arithmetic mean of known class embeddings or a single generic text prompt. However, a single mean proxy compresses diverse semantic directions into one representation and may fall into semantically ambiguous regions, a limitation we term mean collapse, thereby restricting its ability to capture multimodal unknown semantics. To overcome this geometric flaw, we propose the Manifold-Guided Multi-Prototype Expansion (MGME) framework. Instead of compressing unknowns into a singular proxy, MGME models them as a diverse set of unknown prototypes distributed along the valid boundaries of known classes via orthogonal geodesic traversal. Specifically, we extract a focused orthogonal residual subspace to serve as a collision-free exploration space for novelty generation. Furthermore, we design a dual-constraint selection mechanism that combines cross-class boundary pruning with large language model-driven semantic steering to ensure the generated prototypes are both topologically safe and semantically aligned. Finally, we introduce Discriminative Prototype Orthogonalization during the training phase to push known text prototypes apart, establishing highly discriminative decision boundaries. Extensive experiments across five real-world object detection datasets demonstrate that MGME effectively prevents manifold collisions and generates novel unknown prototypes with strong visual objectness, achieving state-of-the-art performance and yielding significant improvements in unknown object discovery.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.