acceptodds
Under review as a conference paper at ICLR 2027

MLServingBench

Abstract

We introduce MLServingBench, a benchmark for evaluating coding agents at a long horizon by building inference serving systems. Given a model architecture, compute topology, and workload, an agent has 12 to 36 hours in a sandbox with up to 8 H200 GPUs to build a correct, performant system. Our tasks span text generation, agentic workload, and speech and multimodal models, using production workloads. Unlike prior benchmarks that focus on localized code changes or tuning an existing engine, MLServingBench requires end-to-end system building across kernels, memory management, scheduling, and parallelism. To evaluate correctness, we introduce a token-level verifier for text model serving that compares outputs from the agent’s solution against a reference implementation, enabling low-cost and high-fidelity verification. We score correct submissions by their speedup over a production serving framework. We evaluat frontier proprietary and open models across seven tasks and find that no agent outperforms the production baseline on every task, with only 23.8% of runs producing a correct system faster than the baseline. We aim for MLServingBench to guide the development of agents that can reliably build efficient serving systems for new model architectures and workloads.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.