acceptodds
Under review as a conference paper at ICLR 2027

Vision Harnessing Agent for Open Ad-hoc Segmentation

Abstract

Segmentation has become easy when the concept is known, requiring retrieval of a learned visual grounding from text. It remains hard for open ad-hoc concepts, where the grounding may not exist as one learned mask and must often be constructed from image evidence through parts, relations, exclusions, and collections. We propose a Vision-guided Ad-hoc Segmentation Agent (VASA), the first vision harnessing agent for open ad-hoc segmentation. VASA is training-free and couples a VLM agent, a segmentation foundation model, and a visual harness that maintains a working mask to make visual progress persistent, inspectable, and editable. Rather than revising text prompts alone, it plans visual operations, invokes segmentation tools, inspects results, edits the mask, and recovers from errors. We construct PARS, a new benchmark that turns part-level labels into open ad-hoc concepts through long-form definition queries. We show that VASA is consistently effective across six VLMs with varying capabilities. Using Qwen3-VL 32B Thinking as the VLM, VASA outperforms various baselines on PARS, surpassing SAM3 Agent by 13.5%–25.3%. On RefCOCOm, VASA improves over SAM3 Agent by 4.8%–8.8% and over other agentic baselines by more. VASA also remains competitive with SAM3 Agent on ReasonSeg for common, named concepts. These results validate VASA's agentic visual construction for open ad-hoc segmentation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.