acceptodds
Under review as a conference paper at ICLR 2027

PerfDiag: An Execution-Free Benchmark for Diagnosing Performance Regressions in AI Infrastructure

Abstract

A performance regression is a commit that changes no output and makes the system worse to run: throughput down, memory up, a tail latency that doubled. In a serving or training stack it has no failing test, and reproducing it needs hardware most researchers cannot rent. The fix is small; finding it is the work. So performance benchmarks ask a tractable question instead: make this code faster, and we will time it. We ask the diagnostic one: given the commit that introduced a regression, name the mechanism from a fixed 32-label set and cite the code site that has to change; the human fix commit is the answer key, and nothing is executed. We release PerfDiag: 430 (regression-inducing, fixing) commit pairs from four production AI-infrastructure repositories, with a read-only harness, decoy negatives for bounding false positives, and a measured zero-call baseline on cause, mechanism and fix site. The task is hard. The best of six models names the cause family on about 58% of items, and the two remedies the task invites do not separate: no prompt scaffold among twenty separates from a single flat call on the one model we swept, and a model-authored three-layer diagnostic decision graph walked one layer at a time is worse than one on three of the five models we could run it on and on neither of the two weakest. §4 reports the resolution of that sweep: 10.2 to 17.9 accuracy points. The mechanism keys on 94% of the scored items were written by Claude Opus 5, which ranks first, so we measure how reproducible they are and not that they are correct. A key it had no part in writing leaves the ranking unchanged. Every number is on train and dev; no run has scored the test split, though an author saw the gold label of 11 of its items during the label audit (Appendix I).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.