Light SAM3 Agent: Efficient Agentic Segmentation via Latent NPs Extraction and Action Space Slimming
Abstract
SAM3 Agent has demonstrated strong performance in reasoning segmentation tasks by combining multimodal reasoning with SAM3's precise mask generation. However, its substantial token consumption and long inference time limit its use in real-world applications. In this paper, we analyze the interaction behavior of SAM3 Agent and identify two key sources of redundancy in its original design: i) agent loop redundancy: SAM3 Agent relies on effective noun phrases (NPs) to obtain valid masks from SAM3, yet in challenging cases, identifying such an NP can consume many or even all available interaction rounds, leading to excessive MLLM inference cost. ii) action space redundancy: SAM3 Agent defines four executable actions, but we find that some contribute very little and may even have negative effects due to the additional context they introduce. To address these issues, we propose Light SAM3 Agent, a training-free acceleration framework that improves token efficiency by optimizing the agent loop and action space. Specifically, for agent loop redundancy, we propose Latent NPs Extraction, which explores alternative NPs in the latent space via probabilistic sampling, thereby compressing multiple rounds of NP exploration into a single round while avoiding costly explicit token generation. Moreover, for action space redundancy, we propose Action Space Slimming, which evaluates the invocation frequency and empirical utility of each action, and removes low-value actions from the candidate action set. Light SAM3 Agent achieves an average end-to-end speedup of 2.11× on the ReasonSeg test set while reducing token consumption by 38.8%. Meanwhile, with a cleaner reasoning context and more effective NP exploration, Light SAM3 Agent improves gIoU and cIoU over SAM3 Agent by 6.2% and 3.4%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.