HOW FEEDBACK SHAPES SEARCH: Candidate Competition in LLM Agents
Abstract
Large language model (LLM) agents use feedback to guide search, but better proposals need not yield a better final solution. We study two LLMs searching fixed sets of image and text classifiers trained before search. With evaluation and selection fixed, showing the correct condition label for each score improves average proposal quality for the requested condition in all four original model–task settings. Paired post hoc comparisons support larger effects on proposal quality than on final quality in two settings, with limited confirmation from new responses. In text classification, the same comparison format encourages one model to revisit tried configurations or explore untried ones depending on score visibility. We construct searches with identically distributed complete proposal-quality sequences but different final outcomes under one selection rule, and characterize when improving a final proposal cannot reduce selected quality. Evaluating feedback therefore requires tracking candidate competition and improvements preserved through selection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.