LeetDecoding: An Experimental Survey and Scalability Evaluation for Prefilling With Causal Linear Attention
Abstract
Significant progress has been made in accelerating transformer-based large language models (LLMs). Among various promising approaches, *exponentially decaying causal linear attention* has emerged as a highly effective operator, yet its existing implementations remain fragmented across the literature. In this paper, we present an experimental survey of this fundamental operator for LLM prefill inference. We also introduce `LeetDecoding`, the first comprehensive library and benchmark providing extensive computation routines and a unified testbed for scalability evaluation. `LeetDecoding` addresses multiple critical gaps in the field: the absence of rigorous complexity analyses, the lack of a systematic consolidation of existing computation methods (previously scattered as secondary contributions without direct comparison), and the dearth of extensive empirical evaluations. Our findings reveal that no single computation method universally dominates across diverse configurations, and that hardware-aware optimizations are paramount for substantially improving overall efficiency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.