Rethinking Over-Segmentation in Few-Shot Object Detection with Vision Foundation Models: A Training-Free Prototype Refinement and Occlusion-Aware Discrimination Approach
Abstract
Few-shot object detection pipelines built on frozen vision foundation models achieve strong cross-domain performance without task-specific training, yet still suffer from over-segmentation: a single object is fragmented into multiple proposals, and non-maximum suppression can discard the correct whole-object proposal in favour of a higher-confidence fragment. Our diagnostic study shows that this failure mode differs qualitatively from classical over-segmentation: proposal coverage is rarely the bottleneck; instead, classification confidence becomes unreliable in the few-shot regime, where biased support sets shift the class prototype toward part-views. Building on this insight, we propose two training-free, complementary modules. Training-free Prototype Refinement analytically performs a one-step contrastive prototype update against an adjacency set of negative samples, bounded by a trust region derived from the support set. Occlusion-Aware Discrimination distinguishes whole-object proposals from fragments by their confidence decline rate under partial occlusion, and additionally retains genuine small objects that NMS would otherwise delete. On eight FSOD benchmarks, our method consistently improves over VFM-based detectors while adding no trainable parameters and only modest inference overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.