acceptodds
Under review as a conference paper at ICLR 2027

Mirage Discoveries: Separating Signal From Selection Noise in LLM-Driven Program Search

Abstract

Autonomous discovery systems increasingly pair language models with evaluators to generate, test, and select new programs end to end. When the same evaluation regime both guides search and decides acceptance, the search optimizes against its own evaluator: as more candidates are tried, those selected increasingly exploit noise in the evaluation slice rather than transfer to unseen data. We measure this winner's curse in LLM-guided evolution of learning-to-rank pipelines over noisy logged user behavior. Fitness-based acceptance identifies the held-out best candidate in 5 of 12 runs, and equal-budget random search transfers as well as guided search. We introduce audited acceptance, which separates search from deployment: candidates are optimized on a cheap fitness fold but enter the deployable set only when their improvement persists on an audit fold, disjoint from both the fitness fold and the final test fold, by a margin calibrated to the paired-bootstrap noise floor of the search evaluator, with all audit usage counted and reported. The same criterion yields a pre-search headroom gate that determines whether a search space contains enough measurable signal to justify autonomous optimization at all. In a search space that passes the gate, all three discovered executable pipelines outperform an Optuna-tuned LambdaMART baseline by up to 0.0057 NDCG@10 on roughly 60k untouched test queries, and the discovered programs are readable code that recovers the structure of the original challenge winners. As autonomous experimentation scales, acceptance, not candidate generation, becomes the binding constraint on what these systems can reliably promote.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.