acceptodds
Under review as a conference paper at ICLR 2027

Beyond Representation: Self-supervised Image Encoders to Segment Anything

Abstract

The Segment Anything Model (SAM) has been the state-of-the-art interactive image segmentation paradigm, which enables users to extract target objects with simple prompts. Despite its generalizability to unseen objects, it relies on large-scale annotated images for training. Meanwhile, Self-Supervised Learning (SSL) has yielded image encoders with strong visual representations. We therefore ask whether self-supervised image encoders alone can support “segment anything” without fully supervised training. We present SSL-SAM, a training-free framework based on on-the-shelf self-supervised vision encoders for interactive segmentation from points, boxes, masks, and scribbles. SSL-SAM converts each type of prompt into a support set of foreground and background respectively, and segments the target object by prototype-consistency scoring in the feature space of a frozen encoder. To deal with very sparse prompt (e.g., a single foreground click), we also propose an automatic support expansion strategy to achieve high performance while reducing the user's effort. Across 20 datasets with 5000+ classes, SSL-SAM outperforms the SAM family (SAM, SAM2, and SAM3) across multiple prompt modes, and achieves a 13.49% gain in mean Intersection over Union (mIoU) in box-only mode. These results show that self-supervised vision encoders can be adapted without retraining for promptable segmentation, which avoids the high cost of large-scale annotation and supervised training in the SAM models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.