HaLA: Householder-adapted Linear Attention for Vision Transformers
Abstract
Linear attention offers an efficient alternative to softmax attention, but its representational gap remains poorly understood. We show that part of this gap arises from a broken geometric symmetry. Softmax attention is coordinate-free: rotating queries and keys together preserves their inner products and leaves attention unchanged. ReLU-style linear attention, however, applies a non-negative gate along fixed feature axes, making attention sensitive to the coordinate system. This gate can passively truncate useful semantic directions before aggregation, leading to a failure mode we call coordinate-lazy truncation. To address this, we propose Householder-adapted Linear Attention (HaLA), a geometry-aware linear attention framework that turns passive truncation into active reorientation. By leveraging learnable Householder reflections, HaLA rotates the query-key coordinate system into the non-negative orthant before non-negative projection to maximize information retention. Built on this principle, HaLA introduces a cohesive multi-scale design: block-wise isometries enhance local discriminability, variance-aware modulation stabilizes long-context activation diversity, and cross-head reflections integrate dispersed subspaces through global covariance mixing. Across various vision tasks including long-sequence super-resolution and image generation, HaLA consistently strengthens linear attention and closes much of the gap to softmax attention under efficient linear-time computation. Source code can be found in the supplementary materials.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.