MLServingBench
Abstract
We introduce MLServingBench, a benchmark for evaluating coding agents at a long horizon by building inference serving systems. Given a model architecture, compute topology, and workload, an agent has 12 to 36 hours in a sandbox with up to 8 H200 GPUs to build a correct, performant system. Our tasks span text generation, agentic workload, and speech and multimodal models, using production workloads. Unlike prior benchmarks that focus on localized code changes or tuning an existing engine, MLServingBench requires end-to-end system building across kernels, memory management, scheduling, and parallelism. To evaluate correctness, we introduce a token-level verifier for text model serving that compares outputs from the agent’s solution against a reference implementation, enabling low-cost and high-fidelity verification. We score correct submissions by their speedup over a production serving framework. We evaluat frontier proprietary and open models across seven tasks and find that no agent outperforms the production baseline on every task, with only 23.8% of runs producing a correct system faster than the baseline. We aim for MLServingBench to guide the development of agents that can reliably build efficient serving systems for new model architectures and workloads.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.