acceptodds
Under review as a conference paper at ICLR 2027

Flow-Controlled Scheduling for LLM Inference with Provable Stability Guarantees

Abstract

Large language models (LLMs) have been widely adopted due to their great performance across a wide range of applications. ChatGPT and Gemini now serve hundreds of millions of active users and handle billions of user requests per day, which puts optimizing LLM inference into the spotlight. A key challenge in LLM inference is that decode lengths are unknown. The memory usage for each request grows with generated tokens, which may lead to overflow and cause system instability. To address this concern, we propose a simple flow-control framework that controls the rate at which prompts join the active set. For known output lengths, we prove that, for every arrival rate that is strictly inside the stability region, there exists a periodic policy that can be stable. We derive a necessary condition that any stable system must satisfy and establish sufficient conditions under which our algorithm can provably achieve stability. Experiments show that, compared to commonly used strategies in practice, our approach achieves higher request throughput and lower average latency.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.