acceptodds
Under review as a conference paper at ICLR 2027

AgentServeSim: Simulation-Driven Policy Design for Agent Serving

Abstract

Serving large language model (LLM) agents requires coordinating cache man- agement, routing, and scheduling across model calls separated by tool execution. Simulating these policies is difficult because their decisions change subsequent call arrivals, retained context can be shared across programs, and cache reuse depends on both placement and resource availability. AgentServeSim addresses these cou- pled requirements through a program-level execution model that preserves history and cache state across LLM calls and enforces policy actions within common resource constraints. It derives arrivals from simulated completions, accounts for shared prefixes without duplicating their memory cost, and checks cache availabil- ity at admission. Across 20 comparisons with real serving, AgentServeSim predicts mean program job completion time (JCT) with 4.74% mean absolute percentage er- ror. It also identifies the same winning policy as real serving among six candidates in all three evaluated settings. Using simulation to evaluate LLM-generated KV- management and scheduling code, we discover a policy that achieves 1.64–2.04× speedup in real serving over a default baseline using first-come-first-served (FCFS) scheduling and least-recently-used (LRU) KV cache eviction, and 1.25–1.45× over the strongest evaluated alternative in each setting.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.