acceptodds
Under review as a conference paper at ICLR 2027

ServingBench: Can Agents Deliver Day-0 Model Support in Production Inference Engines?

Abstract

Serving engines such as vLLM now provide day-0 model support, making major open-weight models available on the day they are released. However, supporting a new model correctly and efficiently on day 0 can be non-trivial. Engineers must implement the new model architecture, integrate it with the serving runtime, and, for some models, optimize kernels and backend code for the target accelerator. We therefore ask: can coding agents deliver correct and efficient day-0 model support? To answer this question, we introduce ServingBench, a benchmark of 148 day-0 model support tasks across 10 accelerator tiers from NVIDIA, AMD, and Google. An automated pipeline constructs executable environments from selected model-integration pull requests in vLLM, SGLang, and TensorRT-LLM. Each environment includes correctness checks and an expert performance baseline validated on the target accelerator. Evaluating seven state-of-the-art models, we find that agents resolve 56.7% of tasks on average, but 33% of resolved patches serve more slowly than the engine maintainers’ implementation, by a median of 18%. Even the strongest model in our evaluation (i.e., Claude Opus 5) delivers a correct patch at the expert’s serving speed on only 52.7% of tasks, and 32% of its resolved patches are slower, some at about half the expert’s speed. The pipeline can incorporate new model releases, allowing ServingBench to track progress on a continually evolving systems engineering task.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.