acceptodds
Under review as a conference paper at ICLR 2027

Coarse KV Residency: Decoupling KV Reachability from GPU Memory in Long-Context RL Rollouts

Abstract

Agentic RL rollouts alternate generation with tool waits, during which each sequence's complete KV state must remain in GPU memory. In a replay of a production turn structure, 85.6% of resident KV belongs to sequences idle in such waits. However, this footprint cannot be reduced by eviction or approximation: rollouts are the training data of RL post-training, and perturbing the sparse top-k selection changes the policy being optimized. We propose coarse KV residency, which decouples logical KV reachability from physical residency: the host stores every sequence's complete KV state, while the GPU maintains a bounded pool per sequence, layer and KV group. Pool capacities are sized at admission from measured sparse-attention demand and maintained by warm-start initialization, LRU retention and periodic reallocation. Residency never changes the model's own top-k selection, and missing selected entries are fetched before attention reads them. On Qwen3.5-4B, pooled decoding is bit-exact against full residency and the selected configuration passes all sixteen registered noninferiority endpoints, while allocation remains within 2.5-13% of an oracle, admission seeding reduces K/V traffic by 25.7%, and periodic reallocation reduces replacements by 40.3% on long rollouts. The freed residency yields up to 88% higher aggregate decode throughput through additional concurrent rollouts.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.