acceptodds
Under review as a conference paper at ICLR 2027

LeetDecoding: An Experimental Survey and Scalability Evaluation for Prefilling With Causal Linear Attention

Abstract

Significant progress has been made in accelerating transformer-based large language models (LLMs). Among various promising approaches, *exponentially decaying causal linear attention* has emerged as a highly effective operator, yet its existing implementations remain fragmented across the literature. In this paper, we present an experimental survey of this fundamental operator for LLM prefill inference. We also introduce `LeetDecoding`, the first comprehensive library and benchmark providing extensive computation routines and a unified testbed for scalability evaluation. `LeetDecoding` addresses multiple critical gaps in the field: the absence of rigorous complexity analyses, the lack of a systematic consolidation of existing computation methods (previously scattered as secondary contributions without direct comparison), and the dearth of extensive empirical evaluations. Our findings reveal that no single computation method universally dominates across diverse configurations, and that hardware-aware optimizations are paramount for substantially improving overall efficiency.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.