acceptodds
Under review as a conference paper at ICLR 2027

FUSED BUT NOT RELATIVE: LOW-RANK ATTENTION BIASES SILENTLY DROP TRANSLATION INVARIANCE

Abstract

A rank- additive attention bias written as , with , can be folded into the queries and keys, allowing compatible unmodified FlashAttention kernels to compute biased attention without materializing the bias. Alternative fused solutions require bias-specific kernel support or compilation. Low-rank factorization supports fused training but does not guarantee relativity: may depend on rather than only on . % Prior work proposes learning these factors from initialization but does not evaluate this variant; its reported instantiations use closed forms, factorize a pretrained table, or fit an existing bias, thereby enforcing or inheriting relativity. When trained freely in vision, only of the learned bias energy remains displacement-dependent, compared with for the unfactorized relative bias it replaces. % We restore relativity by generating factors from learned frequencies, whose inner product reduces to and therefore depends only on displacement by algebraic identity. A learned phase lifts the resulting even-symmetry restriction, still depending only on displacement, without increasing factor width. Kernel and computational cost are unchanged, while positional parameters become independent of window size and context length; in Swin-T at window , they fall from to . % At matched factor width, architecture, and training recipe, restoring relativity reduces language-model perplexity degradation at the training context by roughly across ranks –, while using fewer positional parameters and improving in-distribution perplexity. In vision, it improves transfer to semantic segmentation and detection, and out-of-domain generalization at the larger window, while the phase variant strengthens robustness at both. These results provide a practical route to learned relative biases within fused attention across vision, language, and other position-structured domains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.