PACE: Measuring LLMs in Performance-Critical Code Generation under Agentic Workflows
Abstract
Code agents can edit files, invoke compilers, run tests, and revise their implementations. Previous works have evaluated speedup over reference implementations. This can surely help people understand how good a piece of code is, but it has a gap between the speedup number and hardware potential as it depends heavily on the quality of the reference for the workload. To further study resource utilization and understand the code quality generated by agents, a shift from workload-centered view to resource-centered view could help build an intuitive view. Therefore, we introduce PACE (erformance-critical gentic ode valuation), a CPU evaluation framework, with a metric called ffective alibrated ttainment (ECA) built on measuring computing resource units. It measures how closely one bounded agent run approaches a calibrated performance reference for each task. We report a five-model baseline on x86-64 and AArch64 containing 3,000 runs. We also report a 135-run x86-64 pilot with nine conditions and DeepSeek-V4.1-Flash. From these assessments, we find that high pass rates can hide substantial performance gaps and that model rankings change across task tracks and hardware platforms. And a larger interaction budget alone may not improve the delivered code. The PACE repository can be accessed at https://anonymous.4open.science/r/PACE-411F
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.