HOFA: Channel-Wise Hybrid Attention via Outlier Factorization
Abstract
KV cache memory remains a primary bottleneck during long-context autoregressive decoding. While prior hybrid architectures mitigate this by alternating exact and linear attention across macro structural axes (depth, width, or sequence), they impose a coarse tradeoff between decoding efficiency and fine-grained associative recall. We introduce Hybrid Outlier-Factorized Attention (HOFA), a channel-wise intra-head decomposition. Inside every attention head, HOFA statically allocates a small -dimensional outlier subspace to exact causal Softmax attention while routing the remaining inlier channels to a recurrent Gated Linear Attention (GLA) state. Hardware profiling demonstrates that this intra-head split reduces the decoding KV-cache footprint by 43.7% and yields up to a +75.9% peak decoding throughput speedup over standard MHA. Pretraining across 70M, 125M, and 350M parameter scales establishes that HOFA matches MHA perplexity and achieves competitive zero-shot reasoning accuracy. We show that this factorization naturally concentrates global associative matching into the outlier channels while the linear recurrence absorbs background context, with retrieval limits closely following a proposed theoretical extreme-value capacity model.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.