HPCAgent-Bench: A Framework for Sound Benchmarking of Code-Optimizing Agents, Harnesses, and Skills
Abstract
Coding agents are increasingly used to optimize code, but existing benchmarks fix one accelerator, programming model or harness, so they Coding agents are increasingly used to optimize code, but existing benchmarks fix one accelerator, programming model or harness, so they cannot isolate what each component contributes or show whether added skills and tools repay their cost. We present HPCAgent-Bench with 689 kernels in three tracks: loop-level optimization challenges, kernels extracted from HPC codes, and machine-learning models. The framework specifies each kernel once in NumPy, derives C and Fortran references, accepts submissions through a common C ABI or Python interface, and runs them on CPUs or GPUs, on one or multiple nodes. An external judge checks correctness on hidden fuzzed inputs and times each submission, closing paths agents use to exploit the testing system. We introduce two metrics for reporting results: intervention efficacy measures an intervention, such as adding a tool, through paired runs on speedup, success rate and cost, and weighted token cost reports input, cached-input and output tokens under explicit weights. Agents reach up to speedup over automatic parallelizers, and their optimizations transfer to another machine: – of answers without device-specific code remain correct, and per-kernel speedups keep a rank correlation of . Across 40 matched setups, interventions significantly change speedup only twice but token cost by up to . Because token weighting decides whether a saving is significant, we state the weights and release every token count as a standard for cost reporting. We provide the first comprehensive, broad and hardened benchmark for the sound evaluation of agentic optimization, with transparent cost metrics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.