Feedback Alignment in Self-Improving Agents
Abstract
Self-improving agents often use the same evaluation signal both to select updates and to report progress. When that signal captures only part of the intended objective, a rising score does not necessarily mean better task performance: the agent may instead be getting better at satisfying the evaluator. We study this failure mode in scaffold-level self-improvement, where a fixed language model repeatedly revises the executable program used to solve future tasks. We introduce a controlled protocol that separates visible feedback used for scaffold selection from held-out target evaluation used only for reporting. Across two-voice counterpoint repair and cross-database SQL repair, scaffold evolution can raise the visible score while lowering the independently measured target. Ordinary best-of-five selection also reduces the counterpoint target, but less than the full evolution procedure under the same call cap. Retained scaffolds further encode evaluator-specific prompting, checking, and selection rules. These results indicate that a rising optimization score is insufficient evidence of improvement; in the tested comparisons, feedback alignment predicts the direction of target change.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.