ATTNSEG: Attention as the Interface Between Language Models and Segmenters
Abstract
Vision-language models for segmentation must communicate where a referred object lies to a pixel decoder. Existing approaches do so by either serializing the mask into the token stream or compressing spatial information into the hidden state of a single [SEG] token. We argue that both overlook a spatial representation the model has already computed: the [SEG] token's attention over image tokens, which is natively defined on the image patch grid and can directly encode the referred region. Probing a frozen segmentation VLM shows that attention contains substantially more directly readable localization information than the [SEG] hidden state at matched capacity. This signal captures object shape rather than merely pointing to the referent and is concentrated in a small subset of late attention heads, while standard uniform aggregation obscures much of this spatial structure. Motivated by these findings, we introduce ATTNSEG, an attention-based interface that makes the VLM the primary localizer and leaves the segmentation decoder to refine its prediction. ATTNSEG learns to aggregate the [SEG] token's attention maps into a dense spatial prompt and directly supervises it with an attention-guidance objective. A two-stage training strategy further enforces this division of labor by first learning localization in the VLM and then training the decoder for mask refinement. Applied to Sa2VA, ATTNSEG improves ReasonSeg by gIoU and GS-Eval by at 4B, and by and at 14B, respectively. Gains reach on object parts, multiple objects, and stuff regions, where identifying where to segment rather than refining boundaries is the primary challenge.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.