acceptodds
Under review as a conference paper at ICLR 2027

ART: Attention Run-time Termination for Efficient Large Language Model Decoding

Abstract

Long-context decoding in Large Language Models (LLMs) is constrained by the cost of accessing and processing the Key-Value (KV) cache. Despite evidence that attention outputs depend jointly on keys and values, most existing KV management methods rely on key-only pruning, since incorporating values incurs prohibitive overhead. In this paper, we propose Attention Run-time Termination (ART), a lightweight run-time mechanism that tracks accumulated attention outputs during kernel execution and terminates subsequent KV block accesses when successive observed output updates remain small. Rather than replacing KV selection, ART dynamically terminates redundant KV traversal on top of existing dense or sparse attention policies. We introduce a stability-based criterion that monitors both magnitude and directional changes of intermediate attention outputs and provides a theoretical characterization of the resulting truncation error. Experiments on LongBench show that ART increases the generation throughput of existing KV-cache methods by up to 20% with marginal quality loss. Even under the demanding distant-retrieval conditions of RULER NIAH, ART provides additional efficiency gains to importance-based KV-cache methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.