Progressive Conformal Prediction for Remaining-Length Estimation in Large Language Models
Abstract
Autoregressive decoding creates an “unknown horizon”: although serving engines allocate KV pages on demand, admission and preemption decisions are made before a request's future demand is known. We introduce Progressive Mondrian Conformal Prediction, a training-free method that converts stopping probability, predictive entropy, and decoding progress into online upper bounds on remaining length. We distinguish an empirical pooled estimator from prompt-exchangeable fixed-step and two-fold request-level constructions; the latter provide finite-sample step or whole-request coverage without assuming independence among states of one request. Across models from 3.8B to 70B, the pooled method reduces mean width by 41–49% relative to global calibration at the same nominal level. The request construction records 4.8% empirical request violations at a 10% budget. A vLLM prototype confirms that the bound can drive admission metadata, while the statistical theorem is explicitly separated from empirical multi-request OOM and preemption outcomes. The bound adds 0.22% per-state overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.