AbridgeKV: Compressing Long Reasoning into the KV Cache of Learned Pause Tokens
Abstract
Long chain-of-thought decoding is memory-bound: the KV cache of the generated tokens grows linearly with the test-time compute being spent. Evicting cache entries does not fix this, because a model never trained on a cache with entries removed generates longer, often up to the length limit. We instead train the model to compress its own output. writes each completed block of generated text into the KV entries of a few learned pause tokens and frees the block's cache, trained end to end with the next-token loss. Three novel choices make this work: the summary rows take positions inside the block they replace, the row budget is spread across layers by each layer's measured value, and the blocks are large (512 tokens). On 446 competition problems with 32K-token generations, keeping a quarter of the generated cache matches an identically trained full-KV model in accuracy and output length; the strongest published evictor at a matched budget writes far longer, and the closest trained compressor, retrained on the same data, loses 10 points. Served in vLLM, the cache is smaller per sequence and a saturated H100 finishes a problem faster than with full KV.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.