Goodhart’s Law in Performance Benchmarks: Measuring and Mitigating Software Overoptimization
Abstract
Coding agents have shown growing promise for optimizing the performance of real-world software, motivating recent benchmarks that assess this capability through executable workloads. However, these benchmarks typically provide only one or a few visible workloads per task due to the substantial expert effort required to construct high-quality workloads. These workloads then act as both optimization feedback and the final evaluation. Agents can thus obtain large reported speedups by exploiting properties of the observed inputs rather than resolving the underlying performance bottleneck, producing patches that fail to generalize to other valid workloads. To systematically measure this workload overoptimization, we introduce WorkloadPro, an evaluation framework that expands existing benchmarks with diverse, held-out workloads while ensuring that each workload exercises the intended optimization. Measuring against these workloads shows that existing top-rankers substantially overstate performance improvements, and their rankings can reverse. To mitigate this problem, we introduce APO (Agentic Performance Optimization), a profile-guided agentic optimization framework that generates diverse workloads at test time and profiles them to explore diverse optimization targets against these workloads. Our preliminary evaluation using WorkloadPro shows that APO outperforms the top-ranking agents on SWE-fficiency, using the same budget and model, and produces fewer overoptimized patches. These findings show that workload diversity is important both for reliably evaluating coding agents and for guiding them toward performance optimizations that generalize.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.