Harnessing Smoothness: Asymmetric Differential Linear Attention for Vision Transformers
Abstract
Linear attention reduces the quadratic cost of softmax attention, but its inherent smoothing effect often causes redundant information to obscure critical visual cues in the fused representation, leading to performance degradation. Rather than merely treating this smoothness as a limitation, we argue that it is inherently well-suited to modeling visual redundancy, as smooth attention weights naturally capture broadly distributed and redundant patterns. Based on this insight, we propose Asymmetric Differential Linear Attention (ADLA), built upon an asymmetric dual-branch architecture. Specifically, the saliency branch extracts discriminative features via group-weighted key-value summarization, while the redundancy branch provides an explicit prior of visual redundancy via vanilla linear attention. Through their differential interaction, ADLA suppresses visual redundancy while reinforcing critical details, sharpening the attention focus. We further incorporate a dynamic kernel routing mechanism to refine feature focusing. Extensive experiments across diverse vision tasks and ablation studies demonstrate that ADLA consistently achieves strong performance while preserving linear complexity.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.