DINO-KNN: Training-Free Multiclass In-Context Segmentation via Nearest Neighbors
Abstract
In-context segmentation (ICS) segments a target image using one or more annotated reference images. Existing approaches often require task-specific training, combine multiple pretrained models, or rely on multi-stage inference pipelines. Several methods predict one foreground concept at a time. Applying these methods to multiclass segmentation requires combining independently predicted masks, which may assign conflicting labels to the same pixel. We ask whether direct classification of frozen visual features is sufficient for accurate multiclass ICS. We introduce DINO-KNN, a training-free approach that applies \(k\)-NN classification to frozen DINOv3 features without a learned segmentation decoder. Given a single reference image annotated with one or more foreground concepts, DINO-KNN classifies target patches using labeled reference features. Given classes compete within a single multiclass decision, producing mutually exclusive predictions without separate per-class segmentation or post-hoc conflict resolution as in state-of-the-art ICS methods. On multiclass episodes from seven datasets spanning natural images, aerial imagery, underwater scenes, and medical images, DINO-KNN achieves the highest average mIoU among the evaluated SOTA methods, surpassing INSID3, the strongest evaluated baseline, by \(6\) mIoU points. DINO-KNN assigns exactly one semantic label to each pixel. For INSID3, the average percentage of predicted foreground pixels assigned to more than one class is \(42%\). These results establish direct \(k\)-NN classification over frozen DINOv3 features as a simple yet powerful method for training-free multiclass ICS. The code and full setup will be made public.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.