acceptodds
Under review as a conference paper at ICLR 2027

What Attention Aggregation Hides: Attention Entropy Routing for Referring Segmentation

Abstract

Language-guided segmentation requires a model to choose visual context that separates a referred region from nearby distractors. Existing methods mainly process Transformer features formed by aggregation weighted by attention, whose outputs do not explicitly record how concentrated the underlying attention distributions are. We introduce Attention Entropy Routing (AER), which treats normalized multiscale attention entropy as a complementary conditioning signal: language generates a bank of channel modulation vectors, local visual features and aligned entropy route among them, and pooled entropy conditions multiscale fusion. An analysis of single-head attention aggregation with fixed value vectors motivates this design, while heads trained on frozen visual features and controlled routing and spatial permutation studies test its effects at the task and spatial levels. Controlled comparisons show gains over visual routing from local entropy and further gains from global conditioning, while evaluations on RefCOCO, RefCOCO+, RefCOCOg, and remote sensing benchmarks with synthetic nonuniform haze cover natural and degraded visual settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.