acceptodds
Under review as a conference paper at ICLR 2027

IdleServe: A System Layer for Idle Resource Utilization in LLM Serving

Abstract

LLM serving systems retain capacity for peak demand, leaving opportunities to prepare useful information between requests. However, background preparation can delay live traffic, and its outputs may be irrelevant or incorrect. We present IdleServe, a serving system that coordinates idle capacity with the preparation of reusable knowledge for future agentic requests. IdleSensor estimates preparation budgets from serving telemetry and controls background admission and cancellation. IdleTaskBuilder schedules preparation tasks, checks generated artifacts on past requests, and uses outcome feedback to guide subsequent allocation. Applications supply task definitions and adapters; IdleServe manages execution and the resulting knowledge. Implemented on vLLM, the system supports procedures, worked examples, and application references without updating model weights. Across four open-weight models and three agentic benchmarks, evaluations using prepared knowledge improve task score by 4.9 points on average, with a maximum gain of 18.5 points. Separate live-server replays of recorded model-call sequences reduce p95 episode latency by 8.7% on average and increase SLO attainment by 1.8 points. Component experiments show that relevant content and verification contribute to quality. Without admission and cancellation, background work increases p95 latency by up to 238% relative to serving without background work; IdleServe's controls keep p95 latency near that baseline. Together, these results support a serving-system approach to turning idle capacity into useful preparation for later requests.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.