acceptodds
Under review as a conference paper at ICLR 2027

PerSeg: Learning Discriminative Prompts for Mask-Free Personalized Segmentation

Abstract

Prompt spaces in foundation models support semantic generalization, but can be insufficient for personalization, where a particular subject must be distinguished from visually similar instances of the same category. In this work, we study mask-free one-shot personalized segmentation: given a single RGB reference image of an unseen subject, without a reference mask or per-identity optimization, the goal is to segment that subject in a query image or video. We introduce , a framework that predicts personalized prompts from visual evidence while explicitly addressing fine-grained negative distractors. First, learns a personalized prompt teacher using positive supervision together with hard negatives, including same-class distractors. Second, amortizes this teacher by predicting an identity-specific offset from a semantic base prompt directly from the RGB reference image, enabling feed-forward personalization of unseen subjects. Third, predicts similarity-gated prompt updates to accommodate target appearance changes over time in video. To systematically evaluate identity discrimination, we introduce the benchmark, a controlled protocol in which the target remains fixed while distractors progress from different-category instances to visually similar and near-identity subjects. On near-identity distractors, negative-aware PerSeg-DPT improves mIoU from to over standard prompt tuning, while mask-free PerSegDelta achieves , approaching the teacher ceiling. PerSegDelta further obtains gains of points over the previous best method on the PODS and DreamBooth composite datasets. On ILIAS personalized retrieval, its prompt embeddings achieve a best mAP of , outperforming class-only and DINOv2 feature baselines. In video, PerSeg-Adapt improves performance by points over PerSegDelta on held-out YouTube-VOS validation. These results show that personalized prompts can be predicted from only a single RGB reference for unseen subjects, and effective personalization requires identity-specific visual representations and explicit supervision against hard negative distractors.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.