acceptodds
Under review as a conference paper at ICLR 2027

MESH: Recursive Self-Improvement through Evolving Search Control

Abstract

Language-model agents increasingly automate discovery by iteratively proposing, implementing, and evaluating candidate solutions. Each execution consumes part of a limited budget and reveals information about the task, and using this information to choose subsequent executions enables recursive self-improvement. An evaluation outcome not only measures the quality of a candidate but also provides evidence about the hypothesis that motivated it. However, existing tree-search agents rely primarily on scores and overlook how outcomes support or challenge the underlying hypotheses. To address this limitation, we propose MESH, a framework in which candidate solutions co-evolve with the explicit beliefs that direct them. At the high level, an agent maintains a global belief state of working hypotheses, each specifying a target candidate, a mechanism grounded in prior evidence, an intervention, and an expected outcome. At the low level, a search process develops candidate solutions. Beliefs compete for the execution budget according to the improvements they have produced. Each selected belief directs a bounded sequence of executions that iteratively apply and refine its intervention, allowing a promising idea to develop beyond a weak initial implementation. The high-level agent periodically revises all beliefs jointly based on the shared search history, so that even failed executions can refine, retarget, or replace hypotheses across the search. The resulting belief-guided recursive loop operates within a single discovery run and is grounded in realized progress rather than the agent’s self-assessment. Experiments on 19 tasks from MLGym and MLE-bench show that MESH attains the best score on 17 of them, including all eight MLE-bench competitions, when compared with linear, tree-structured, and graph-structured search controllers under the same worker models and execution budgets. Further analyses show that belief-directed executions improve the incumbent three times as often as score-driven ones, with the gap widening as score-driven search stalls, and that freezing the beliefs forfeits up to 59% of the gain.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.