acceptodds
Under review as a conference paper at ICLR 2027

Representative Attention for Vision Transformers

Abstract

Many efficient attention methods reduce the quadratic cost of global self-attention by using predefined spatial partitions to construct compact proxies or restrict interactions. However, these partitions can separate distant tokens with similar features and group neighboring tokens with dissimilar features. This creates a mismatch between token interactions and the feature structure of visual content. To address this mismatch, we introduce Representative Attention (RPAttention), which organizes global communication through a compact set of input-dependent representatives. A learned projection of input token features determines their contributions to each representative. RPAttention then constructs representatives by GATHERING information from ALL image tokens rather than a predefined spatial subset. The constructed representatives then INTERACT through lightweight self-attention to refine the gathered content. They DISTRIBUTE their refined content to each spatial query using attention weights computed separately from the gathering assignments. With M representatives and N spatial tokens, the complete Gather–Interact–Distribute process requires O(NM+M^2) token interactions. Experiments on image classification, object detection, instance segmentation, and semantic segmentation demonstrate strong performance across multiple backbone families and competitive results against recent efficient attention methods. Runtime measurements further show favorable scaling as input resolution increases.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.