InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents
Abstract
AI agents are increasingly used to automate research and development, but existing benchmarks often constrain them to prescribed workflows or narrow action spaces, making it difficult to distinguish actual strong performance from reproductions of pre-existing solutions. We introduce InferenceBench, where an agent must deploy an OpenAI-compatible inference server and optimize LLM inference speed. Each agent receives a target LLM, one H100 GPU, an optimization scenario, and a 2-hour time budget. The objectives each address a bottleneck in LLM inference: the delay before the first token, the delay between later tokens, and the number of requests served per second under load, with a fourth objective that balances all three at the same time. We evaluate 42 agent configurations and find that all agents improve over a na\"ive PyTorch server, and 30 of 42 exceed the best serving engine with default settings ( for vLLM). However, a simple hyperparameter search over vLLM hyperparameters with the same time budget reaches , beating all but a single agent, Claude Opus 5.5 (). Although agents enumerate many relevant optimization techniques, they overwhelmingly converge on a single inference framework, test few distinct configurations, and spend the remaining budget re-measuring, repairing, or optimizing hyperparameters rather than exploring substantially different strategies. Overall, InferenceBench measures how well agents operate in an open-ended AI-engineering setting, where memorized solutions give only limited gains and the unbounded speedup score leaves room to track stronger agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.