PerfDiag: An Execution-Free Benchmark for Diagnosing Performance Regressions in AI Infrastructure
Abstract
A performance regression is a commit that changes no output and makes the system worse to run: throughput down, memory up, a tail latency that doubled. In a serving or training stack it has no failing test, and reproducing it needs hardware most researchers cannot rent. The fix is small; finding it is the work. So performance benchmarks ask a tractable question instead: make this code faster, and we will time it. We ask the diagnostic one: given the commit that introduced a regression, name the mechanism from a fixed 32-label set and cite the code site that has to change; the human fix commit is the answer key, and nothing is executed. We release PerfDiag: 430 (regression-inducing, fixing) commit pairs from four production AI-infrastructure repositories, with a read-only harness, decoy negatives for bounding false positives, and a measured zero-call baseline on cause, mechanism and fix site. The task is hard. The best of six models names the cause family on about 58% of items, and the two remedies the task invites do not separate: no prompt scaffold among twenty separates from a single flat call on the one model we swept, and a model-authored three-layer diagnostic decision graph walked one layer at a time is worse than one on three of the five models we could run it on and on neither of the two weakest. §4 reports the resolution of that sweep: 10.2 to 17.9 accuracy points. The mechanism keys on 94% of the scored items were written by Claude Opus 5, which ranks first, so we measure how reproducible they are and not that they are correct. A key it had no part in writing leaves the ranking unchanged. Every number is on train and dev; no run has scored the test split, though an author saw the gold label of 11 of its items during the label audit (Appendix I).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.