When Does Feedback Help? Evaluating Feedback for Iterative Refinement in Coding Agents
Abstract
Recent advances in AI increasingly rely on agents that iteratively refine their responses at test time using feedback from an environment. Yet it remains unclear what type of feedback most effectively enables an agent to make discoveries beyond its existing capabilities. We investigate this question using a controlled testbed of coding problems spanning different difficulty levels. We compare six forms of feedback most commonly used in practice: detailed execution reports, aggregate performance scores, self-reflection, external LLM code review, failures on LLM-generated tests, and ground-truth hints written by human experts. Using this framework, we show that feedback effectiveness depends strongly on problem difficulty: On easy problems, nearly all feedback types enable the agent to reach full accuracy when the agent was unable to solve them in the initial attempt. On medium problems, feedback derived from ground-truth solutions yields the largest improvement (40.4%), followed by feedback from LLM-generated test cases (24.1%). On hard problems, however, even ground-truth hints provide only a marginal improvement (5.3%). In this case, feedback based on LLM-generated test cases remains the most effective (20.4%). Nevertheless, under the same inference budget, best-of-N parallel sampling outperforms every iterative-feedback method in hard problems. In fact, we show that the gains from refinement are proportional to the gap between the initial response and the best solution obtainable through parallel sampling. These findings reveal an illusion of sequential refinement: on difficult problems, iterative refinement primarily helps the agent recover from poor initial generations and approach, but not surpass, the best solutions it could independently generate.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.