LMSysDojo: Building Scalable Analytical Agents for Optimizing Complex LLM Serving Systems
Abstract
Complex real-world computing systems are hard to optimize because their performance emerges from many interacting components and depends on the hardware and workloads they serve. We introduce **LMSysDojo**, a toolkit that reconstructs reproducible optimization playgrounds from merged performance pull requests in vLLM and SGLang. Agents optimize live deployments on real GPUs, and each patch's improvement is normalized against the human expert's fix on the same setup. On the held-out splits of our **LMSysDojo-100** benchmark, a Claude Code agent using Opus 5 achieves –% of the expert's gain. A key bottleneck is that a patch's worth is known only after costly deployment and measurement, and a vanilla coding agent exposes no ex-ante signal for choosing among candidate ideas beforehand. Inspired by how systems researchers justify optimizations, we propose *gap analysis*, which separates diagnosis from implementation: an analyzer locates a bottleneck, proposes a mechanism, and predicts its gain before an implementer builds it. *Parallel gap analysis* selects among competing analyses by predicted gain; *structured gap analysis* expresses each prediction as a program over profiles, experimental evidence, and stated assumptions; and *iterative gap analysis* checks and refines these programs before selection. Our combined approach reaches –% of the expert's gain, improving over the single agent by – percentage points. Across variants, predictive signals turn analysis into a new scaling axis under a fixed evaluation budget, with gains continuing through analyses. We will release the toolkit, benchmark, and -setup corpus.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.