acceptodds
Under review as a conference paper at ICLR 2027

EffiPair: Improving the Efficiency of LLM-generated Code with Differential Execution Feedback

Abstract

Large language models (LLMs) can generate functionally correct programs that differ substantially in execution efficiency. Existing inference-time optimization methods typically refine each candidate using pointwise runtime or profiling feedback, which identifies how costly an implementation is or where the cost arises, but offers limited guidance on how the computation should change. We introduce DIFFERENTIAL EXECUTION FEEDBACK (DEF), which compares implementations that are nearby in structural program space but separated in performance, turning their execution and implementation differences into directional optimization evidence. We demonstrate DEF in EffiPair, a training-free, test-time framework that pairs structurally similar programs with different efficiencies, distills their relative execution behavior into compact feedback, and iteratively refines a candidate pool. Across EvalPerf, Mercury, and ENAMEL, using GPT-4o mini, DeepSeek-V4.1 Flash, and GPT-5 mini, EffiPair achieves the highest value on each benchmark’s official efficiency metric in all nine model–benchmark settings under matched evaluation conditions. Moreover, two contrastive refinement rounds improve efficiency over the selected initial draft in every setting while preserving or improving Pass@1. These results demonstrate the effectiveness of relational execution feedback as a lightweight signal for test-time code optimization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.