acceptodds
Under review as a conference paper at ICLR 2027

Improved Belief-Attention in Vision Tasks

Abstract

Recently, Belief-Attention was proposed to improve Transformer performance by orthogonally projecting the softmax-weighted sum of vectors relative to the original vectors and retaining the perpendicular component as a residual signal. In this paper, we first present an ablation study demonstrating that the projected component also encodes valuable token correlation information that should not be overlooked. Based on this insight, we propose Belief2-Attention, an extension that effectively leverages both components. Specifically, the projected component passes through an activation function and a linear mapping before being merged with the target token—effectively forming an internal two-layer feedforward network (FFN) within the attention mechanism. Furthermore, while standard attention models token correlations solely via the inner product , we modify this formulation to to enrich pairwise token representations for vision tasks. This allows to capture symmetric correlations, leaving to model complementary directional dependencies. We demonstrate that Belief2-Attention possesses greater expressive capacity than standard attention under mild conditions. Finally, we validate the effectiveness of Belief2-Attention on vision tasks, including image classification and image segmentation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.