AccelMark-Agent: Benchmarking Inference-Optimization Agents Across Accelerators and Workloads
Abstract
Which serving plan is fastest depends on the deployment it runs in, so measuring how well an agent finds one requires evaluating across environments, and measuring it reliably requires a throughput number the agent did not produce for itself. We introduce \sysname, a benchmark for inference-optimization agents that supplies both. Five workloads stress different stages of the serving pipeline and run on three accelerators spanning three architectural generations, with the agent's backbone and its harness as two further axes. A trusted controller holds the measurement channel: the agent submits one atomic candidate per round and never measures, every reported gain is the median of three independent re-measurements behind an accuracy gate, failed and no-op rounds stay in every denominator, and delivered artifacts are fingerprinted so an improvement can be told apart from an unchanged baseline. Ten repeats of a fixed runner place the measurement band at . Over recorded rounds the benchmark separates agents on every factor and saturates on none: task medians run from to , backbones from zero to seven successful cells out of eight, and harness preference reverses between backbones on the same task. Because the protocol records what was delivered and not only what it scored, it also makes legible a regularity its axes were not designed to find: the size of what an agent returns tracks the configuration headroom it was handed. Holding the workload, the device and the stage-restriction prompts fixed and moving the starting configuration turns a stage that returns nothing into one that returns . The largest result in the study, , comes from a different mechanism: an agent detecting that the baseline runner had no sliding-window-attention operator and implementing one. An evaluation sampling one workload at one starting point cannot separate these regimes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.