acceptodds
Under review as a conference paper at ICLR 2027

ReasonAlloc: Hierarchical Decoding-Time KV Cache Budget Allocation for Reasoning Models

Abstract

Long reasoning sequences make the key-value (KV) cache a major source of inference memory use and, consequently, a critical inference bottleneck. Existing decoding-time compression methods typically propose to select tokens under uniform budgets, overlooking differences in demand across layers and KV heads. Non-uniform methods primarily target prefill compression and are not adequately tailored to the growing cache demands of reasoning models during decoding. Therefore, how to allocate KV-cache budgets across layers and KV heads according to their decoding-time demand remains largely unexplored. To bridge this gap, we propose ReasonAlloc, a training-free method for hierarchical KV-cache budget allocation during decoding. ReasonAlloc is based on our systematic attention analysis across five models and four datasets, which shows that layer-level cache demand varies across models but is similar across datasets for a given model. Within a layer, demand is concentrated in a subset of query heads. These findings motivate calibrating layer budgets once per model and adapting KV-head budgets to demand during decoding using current token scores. We decouple allocation from token selection, using current token scores to guide head budgets while retaining existing eviction rules. Experiments on mathematical reasoning and code generation show consistent gains over R-KV at the same cache budget. On DeepSeek-R1-Distill-Llama-8B, ReasonAlloc improves MATH-500 accuracy by 4.55-6.45 percentage points at average per-head budgets of 256, 512, and 1,024 tokens. Its throughput remains within 0.66% of R-KV at matched batch sizes and budgets. At 16K generation, cache compression enables larger batches and up to 5.52× the throughput of an uncompressed cache.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.