acceptodds
Under review as a conference paper at ICLR 2027

Prompt Optimization Enables Improving and Self-Distilling Image Segmentation Models

Abstract

Text-Promptable segmentation models (PSMs) like SAM 3 and Falcon Perception can generate masks from natural language. However, we know from other model types like LLMs that prompt optimization can greatly improve outputs, and that these improved outputs can be used to self-train the model for recursive self-improvement. To test this for segmentation models, we study whether PSMs can be improved without ground-truth labels by optimizing the prompts using multimodal LLMs. First, we test and find that training-free iterative prompt refinement, where an MLLM inspects the image and current mask to propose improved prompts, raises segmentation quality by approximately 20 IoU points over category-name prompts. Then, we train MLLMs via GRPO to compress multi-step refinement into a single prompt rewrite, which achieves similar quality at a lower inference cost. Finally, we close the loop by distilling optimized masks back into the segmentation model, so that it produces improved masks directly from the original prompt. Through this work, we enable label-free self-improvement for PSMs for the first time and provide a foundation for autonomous adaptation to new domains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.