Relay, Don’t Route: Adaptive Population Handoff for Cost-Efficient LLM-Driven Evolution
Abstract
Large language model (LLM)-driven evolution has shown promise for program search and algorithm discovery, but relying on strong models throughout long evolutionary runs is costly. A natural alternative is to combine cheap and strong models under a fixed inference budget. However, existing approaches typically allocate models at the level of individual queries or mutation steps, overlooking that evolutionary search is stateful: each generated candidate changes the population from which subsequent mutations are produced. We empirically analyze LLM-driven evolutionary trajectories and find that search progress is strongly front-loaded, early trajectory performance is informative but noisy, and cheap models recover much of the early progress achieved by strong models at lower cost. Motivated by these findings, we propose \model, a training-free framework that shifts budget allocation from individual calls to evolving populations through adaptive population handoff. A cheap model explores multiple trajectories in short blocks allocated by a bandit scheduler. Relay Gain, defined as the marginal improvement of a compact, quality-diverse candidate bank constructed for handoff, serves as the scheduler reward and determines when to hand off. The curated candidates initialize a shared strong model population for refinement. Across four program-evolution benchmarks, \model remains competitive with the strongest baselines despite task-dependent differences in which single model performs best, with its clearest gains on Circle Packing (Square). Our results suggest that in stateful search, budget allocation should be organized around the population, not the individual call.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.