Feedback Is Not a Free Lunch: The Exploration Cost of Auxiliary Information in Agentic Algorithm Discovery
Abstract
LLM agents discover algorithms by writing programs, testing them, and revising them based on the results. A test gives a score and can also report truthful details about a program’s behavior. We call these extra details auxiliary feedback. They may suggest useful changes, but we find that even truthful feedback is not a free lunch. In a matched study of 15 algorithmic and engineering tasks, adding auxiliary feedback raises the best score found across repeated runs on five tasks, lowers it on eight, and leaves two unchanged. To understand these results, we examine the programs agents try in a data-center control task. For each of five models, the best run with sustained auxiliary feedback earns a lower score and tries fewer new ways of using information to choose control actions than the best run without the extra feedback. We call this the exploration cost of feedback. Diagnostics can help an agent improve a promising idea, but repeatedly following them can leave fewer opportunities to try other ideas. We bridge agentic discovery with auxiliary-guided Bayesian optimization, isolating how guidance strength turns useful feedback into restricted exploration. The signal helps when its influence is moderate, but causes search to miss a better region when its influence is too strong. Removing strong guidance later restores exploration. In four tasks where sustained auxiliary feedback lowered the best score, renewed-exploration runs find solutions exceeding the strongest baseline references. Together, these results suggest that feedback should help agents improve promising ideas while leaving room to explore alternatives.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.