acceptodds
Under review as a conference paper at ICLR 2027

TailorAttention: Accelerating LLM Attention with Latency-Aware Chunk Scheduling

Abstract

Modern LLM services handle heterogeneous workloads with sequence lengths varying across tasks and requests. Chunked attention distributes KV-cache chunks across cooperative thread arrays (CTAs), making partitioning and scheduling criti- cal to GPU utilization. Existing kernels, however, pipeline mainly within chunks, incurring boundary underutilization and phase-dependent CTA interference. Con- sequently, equal-sized chunks can have different latencies, making length-based scheduling unreliable and causing load imbalance and tail underutilization. We present TailorAttention, a kernel–scheduler co-design for chunked attention. Its Hopper- and Blackwell-optimized kernels enable inter-chunk pipelining and reduce output-layout conversion through transposed tensor-core dataflows, yielding predictable chunk latencies captured by affine models. Guided by these estimates, a millisecond-scale scheduler jointly selects potentially unequal chunk sizes and CTA assignments while accounting for reduction overhead, with a provable additive bound. Across GPUs, models, and workloads, TailorAttention improves attention- kernel performance by up to 4.96x and reduces end-to-end latency by 27.1% over FlashAttention-3/4 and FlashInfer

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.