acceptodds
Under review as a conference paper at ICLR 2027

InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents

Abstract

AI agents are increasingly used to automate research and development, but existing benchmarks often constrain them to prescribed workflows or narrow action spaces, making it difficult to distinguish actual strong performance from reproductions of pre-existing solutions. We introduce InferenceBench, where an agent must deploy an OpenAI-compatible inference server and optimize LLM inference speed. Each agent receives a target LLM, one H100 GPU, an optimization scenario, and a 2-hour time budget. The objectives each address a bottleneck in LLM inference: the delay before the first token, the delay between later tokens, and the number of requests served per second under load, with a fourth objective that balances all three at the same time. We evaluate 42 agent configurations and find that all agents improve over a na\"ive PyTorch server, and 30 of 42 exceed the best serving engine with default settings ( for vLLM). However, a simple hyperparameter search over vLLM hyperparameters with the same time budget reaches , beating all but a single agent, Claude Opus 5.5 (). Although agents enumerate many relevant optimization techniques, they overwhelmingly converge on a single inference framework, test few distinct configurations, and spend the remaining budget re-measuring, repairing, or optimizing hyperparameters rather than exploring substantially different strategies. Overall, InferenceBench measures how well agents operate in an open-ended AI-engineering setting, where memorized solutions give only limited gains and the unbounded speedup score leaves room to track stronger agents.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.