acceptodds
Under review as a conference paper at ICLR 2027

Better Feedback, Simpler Agents: Rethinking Agentic Code Optimization

Abstract

Modern coding agents are often made more capable by adding specialized agents, tools, memory, and additional reasoning stages. But does adding more machinery actually produce better optimized code? We study this question systematically in code optimization, where an agent must discover faster implementations while keeping the optimization process affordable. We develop a controlled \system with 24 workflows spanning four design choices: direct versus delegated coding, raw versus analyzed execution feedback, optional tools and skills, and how earlier attempts are stored. We evaluate these workflows with GPT-OSS-20B, GPT-OSS-120B, and Qwen3.8-27B on 54 NPBench CPU kernels. The trade-offs are hard to anticipate: more elaborate workflows do not consistently perform better, and most components help in one setting while costing tokens or speed in another. The clearest gain comes from what the main agent receives between attempts: an analysis agent turns execution results into concise guidance, and each turn starts from a compact context. Across both coding modes, analyzed feedback makes programs 1.15-1.46 faster for all three models; with direct coding, the GPT-OSS models gain 1.47-1.51 and use 25-55% fewer tokens, although searches take longer. With analyzed direct coding, tools and skills raise token use by 36-134% without a significant speed-up, and history helps GPT-OSS direct coding more than delegated coding. The configuration with every component enabled uses 54-146% more tokens than direct coding with analysis and history, again without a significant speed-up. On three further models, analyzed feedback helps both Qwen3.5 models and Kimi-K2.7-Code's direct coding, while the value of the other components changes: Qwen3.5-397B, unlike GPT-OSS, gains substantially from delegation. Our results show that code-optimization performance does not grow monotonically with workflow complexity: simpler combinations can produce faster code with fewer tokens.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.