acceptodds
Under review as a conference paper at ICLR 2027

Accelerating Attention Projection with Basis Decomposition

Abstract

Attention is a core operation in large language models (LLMs). We present BD Attention (**BDA**), a *lossless algorithmic reformulation* of attention that accelerates its linear projections. BDA is enabled by a simple matrix identity from Basis Decomposition (**BD**), which restructures multi-head projections into a compact form while preserving attention outputs in exact arithmetic. Complementary to I/O-aware system optimizations such as FlashAttention, BDA provides an architecture-agnostic reduction in projection FLOPs. BDA requires only **4s** of offline preparation on DeepSeek-V2-Lite (16B), **with no retraining required**. On modern GPUs, it achieves **34% faster** key/value projections in FP16 and **25% smaller** K/V projection weights. These gains come with a perplexity (PPL) increase of just **0.02%** (FP16) or **0.0004%** (FP32) on DeepSeek-V2-Lite—a negligible effect on model performance. These results position BDA as a theoretically exact method for lossless attention projection acceleration that is complementary to existing engineering-level optimizations. Our code is available at https://anonymous.4open.science/r/Basis-decomp-57B8.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.