Jumping the Line: Exploiting Length Predictions in LLM Scheduling
Abstract
With the proliferation of large language model (LLM) deployments efficient request scheduling has become increasingly important for reducing completion time and improving user experience. Size-based scheduling policies, such as Shortest Job First, reduce average completion time by prioritizing shorter requests, but require knowledge or estimates of request sizes. In LLM inference, output lengths, and therefore request sizes, are unknown in advance. As a result, LLM schedulers rely on predicted output lengths to guide scheduling decisions. This work proposes JIL, an attack that exposes a vulnerability in such predictionbased LLM schedulers. By manipulating the scheduler’s length prediction, an adversary can cause its requests to receive higher scheduling priority and reduce their completion time. As a case study, we consider TRAIL, a prediction-based LLM scheduling system that uses a lightweight probe to estimate output length. JIL exploits this mechanism by optimizing adversarial suffixes that, when appended to a request, cause the probe to underestimate its output length, thereby increasing the request’s scheduling priority. We evaluate JIL on two datasets and four LLMs across varied request profiles and deployment configurations. Our results show that JIL successfullymanipulates the scheduling signal, reducing predicted output lengths by up to 83.4%. In end-toend serving experiments, adversarial requests complete up to 1.53× faster on average. The reduction in predicted length is substantially larger than the change in actual output length, demonstrating a mismatch between the scheduler’s estimate and the request’s realized size. Response utility varies across models and tasks, sometimes revealing a trade-off between scheduling advantage and response quality. Finally, we evaluate scheduler-side defenses and show that grouping length predictions into coarse intervals reduces JIL’s scheduling advantage and mitigates delays to benign requests.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.