Accelerate Bench: Evaluating Agents for Application-Level Program Optimization
Abstract
Application optimization requires coding agents to discover useful changes and coordinate them across a complete program while preserving the required computation. We introduce Accelerate Bench, a benchmark of 32 CPU optimization tasks adapted from scientific applications and supercomputing competitions. Each task provides a working program and specifies the workload, correctness requirements, timing boundaries, and permitted edits, leaving agents to choose where and how to optimize. Public testing supports iterative development, and a separate verifier scores the final submission. We evaluate 12 coding agents, each pairing a model with a harness. Of their 325 correct outcomes, 161 (49.5%) stay below 1.25× speedup. Controlled replays show why coordination across an application matters: identical source code runs between 0.29× and 2.12× depending only on build and runtime settings, and one code change is worth 1.01× or 6.4× depending on another. Accelerate Bench pairs executable workloads with agents' programs and trajectories to evaluate both their gains and the process behind them.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.