acceptodds
Under review as a conference paper at ICLR 2027

When Code Crawls: Benchmarking Proactive Optimization in Coding Agents

Abstract

On benchmarks such as SWE-bench, coding agents now resolve a substantial share of real software tasks end to end. What those benchmarks score is functional correctness: runtime is not part of the score, so a program that returns the right answer after an hour still passes. Benchmarks that measure runtime either score the efficiency of generated code or ask for a performance improvement outright. In production, speed is part of usefulness, and engineers fix a step that runs far longer than it should instead of letting it finish. We study whether coding agents self-initiate performance optimization when performance is not an explicit task objective. We introduce CrawlBench, which runs the same repair task under two framings: the slowness is either stated in the prompt or left for the agent to discover and repair on its own. Every pathological case is paired with a healthy twin, so efficiency is the only quantity that varies. The reported pool is 192 cases: 88 single-bug Python artifacts, 12 held-out decay tiers, 19 two-bug stacks, 20 regressions mined from merged performance-fix pull requests, and 53 single-bug asymptotic replicas (28 in C, 25 in JavaScript). Across ten production models, agents told the program is slow fix 93.6–97.7% of the single-bug pool; run on the identical task without the hint, the same agents fix 23.5–90.2%. The gap spans +7.2 to +72.0 points. On all ten models, each run for three seeds, the implicit rate has sd ≤4.5 points. Every model repairs 93.6–97.7% once told. Implicit success still requires the agent to discover the slowness, judge it worth fixing, and modify the code unprompted. That self-initiation is what fails. Failure-mode triage finds a miss is the largest failure for every model, at 52–88% of unfixed implicit episodes: the submission stays correct but slow. Fast-and-wrong submissions are 2.0% of those unfixed episodes. We release the benchmark, harness, and all cases.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.