acceptodds
Under review as a conference paper at ICLR 2027

RACE: Calibrated KV-Cache Reservations from Prefill States

Abstract

LLM serving systems make request-admission decisions before the decoding demands of incoming requests are fully known. In prefill–decode disaggregated systems, forecasting output length can help decoding workers plan key–value (KV) cache reservations and determine how many requests to admit. The challenge is that repeated generations from the same prompt can differ in length, making reservations based on a single predicted length vulnerable to unexpectedly long responses. We introduce Rollout-Aware Calibrated Endpoints (RACE), which predicts a token budget by accounting for both typical response length and response-length variability. RACE learns these quantities with two predictors trained on prompt-processing hidden states and multiple sampled responses per training prompt. It assigns larger safety margins to prompts with more variable response lengths, then calibrates these margins on an independent prompt set. On Qwen3-8B and DeepScaleR, RACE reduces the average token budget by 9.94% compared with GBM-Quantile-PAC, a calibrated quantile-regression baseline using the same features. Its budgets accommodate 97.56% of test outputs under the configured output-length limit. These smaller reservations reduce the capacity set aside per request, leaving more room for concurrent work under reservation-based scheduling.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.