AgentSimArena: Benchmarking LLM Serving Simulators for Agent Workloads
Abstract
Serving simulators can reduce the cost of exploring LLM-agent deployments, but inaccurate predictions can lead to poor deployment choices. Evaluating these simulators is difficult because each predicted call completion determines when sub- sequent work arrives, changing future contention. We introduce AgentSimArena, a benchmark that preserves this feedback while comparing simulator predictions with real executions under matched workload and deployment constraints. Across six simulators, eight deployments, and five arrival rates, we evaluate 240 predic- tions against 40 vLLM reference runs. LLMServingSim 2.0 achieves the lowest mean absolute relative error in mean program completion time, at 7.1%; DynoSim achieves 10.6%, with a median simulation runtime of 6.86 s versus 1,224.26 s for LLMServingSim. No simulator remains within 10% error across all settings. In one observed deployment comparison, following a simulator’s recommendation selects an alternative with 6.07× the measured mean program completion time of the other deployment. Arena documents integration extensions, modeling differences, and available preparation-cost evidence, providing a basis for workload-specific simulator selection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.