acceptodds
Under review as a conference paper at ICLR 2027

GABA: Geometry-Aware Bilinear Attention for Linear Transformers

Abstract

Linear attention makes global context tractable in transformers by replacing dense pairwise attention with a compressed state. Recent work has improved what this state encodes through rank augmentation, gating, or multiple local summaries, but the readout remains a coordinate-aligned query-key contraction whose interaction geometry is shared across output channels, limiting output-dependent cross-channel interactions and the attainable representation rank. A second, orthogonal gap concerns fusion: the value representation feeding global aggregation and its locally enhanced counterpart evolve independently and meet only at the output, leaving their relative geometry unconstrained. We introduce Geometry-Aware Bilinear Attention (GABA), which addresses both within a single associative block. An output channel-conditioned low-rank bilinear refinement allows different output channels to realize distinct query–key geometries and contributes complementary state directions. A contrast-guided anti-symmetric interaction displaces the global and local pathways equally and oppositely about their shared midpoint; although their sum is preserved, subsequent branch-specific transformations induce a non-trivial refinement. The all-linear formulation retains linear sequence complexity while keeping model-level FLOPs essentially unchanged. GABA consistently outperforms strong linear-attention baselines across model scales; the hybrid model achieves Top-1 test accuracy on ImageNet-1K, outperforming the second-closest baseline by , as shown in Fig. 1.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.