RSIBench-Infra: Can Agents Optimize Their Own Training and Inference Infrastructure at Production Scale?
Abstract
Improving the efficiency of training and inference infrastructure is a practical component of recursive self-improvement. Existing benchmarks for agent-driven optimization largely focus on individual software layers, leaving unclear whether agents can deliver end-to-end performance gains under production workloads. We introduce RSIBench-Infra, a benchmark of 20 training and inference tasks spanning diverse model families and modalities, with deployments of up to 32 accelerators on NVIDIA GPUs and Ascend NPUs. Agents start from runnable baselines derived from official recipes and can modify the software stack across engine, kernel, and instruction levels. An OCI image-layer submission protocol captures these changes for reproducible deployment, while independent evaluation measures end-to-end performance under task-specific quality constraints. Across six agent–model configurations and 120 trials, 57 trials achieve validated speedups above 1.10×, with the best configuration exceeding this threshold on 60% of tasks. These results demonstrate that current agents can improve production- scale training and inference performance, but identifying effective optimization strategies and consistently translating local improvements into end-to-end gains remain challenging.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.