The Noise Floor of LLM Program Discovery
Abstract
Systems such as FunSearch, AlphaEvolve and ShinkaEvolve put a language model inside a search loop and report gains over plain sampling, usually from one or a few runs. We ask whether those gains are real at the budget most practitioners have, a few hundred LLM calls per problem. Holding the model, the tasks, the call budget and the final selection rule fixed, we compare best-of-N sampling with two island-evolution methods, reflect-then-sample and an EXP3 router on fifteen verified tasks, with 10 to 20 independent seeds per method and task. The result is a noise floor. At 100 calls no method beats sampling by more than 1.03 seed standard deviations on any task, and reflection and adaptive allocation never win; only FunSearch-style evolution has a small, consistent edge. At 300 calls that edge grows large on some instances and turns negative on others, and a pre-registered ladder of covering designs shows that distance from the published record does not predict which. What does move the floor is the model: with a weak 7B model whose first-shot programs are often broken, evolution repairs them and beats sampling clearly, and still ties where the model plateaus. The seed count matters as much as the method. With three seeds per method, the norm in this literature, our own data would have missed four of five effects the full analysis detects and, in one method pair in nine, named a winner it does not support. We therefore report every tie with its minimum detectable effect, every win with its split-half replication rate, and every prediction as written before the data came in, including the ones that failed.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.