acceptodds
Under review as a conference paper at ICLR 2027

AgentSimArena: Benchmarking LLM Serving Simulators for Agent Workloads

Abstract

Serving simulators can reduce the cost of exploring LLM-agent deployments, but inaccurate predictions can lead to poor deployment choices. Evaluating these simulators is difficult because each predicted call completion determines when sub- sequent work arrives, changing future contention. We introduce AgentSimArena, a benchmark that preserves this feedback while comparing simulator predictions with real executions under matched workload and deployment constraints. Across six simulators, eight deployments, and five arrival rates, we evaluate 240 predic- tions against 40 vLLM reference runs. LLMServingSim 2.0 achieves the lowest mean absolute relative error in mean program completion time, at 7.1%; DynoSim achieves 10.6%, with a median simulation runtime of 6.86 s versus 1,224.26 s for LLMServingSim. No simulator remains within 10% error across all settings. In one observed deployment comparison, following a simulator’s recommendation selects an alternative with 6.07× the measured mean program completion time of the other deployment. Arena documents integration extensions, modeling differences, and available preparation-cost evidence, providing a basis for workload-specific simulator selection.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.