Eureka Is a Distribution: Executable Regime Feedback for Language-Agent Algorithm Discovery
Abstract
An algorithm earns its name by surviving beyond the evaluator that shaped it. We call this criterion distributional survival: retaining an advantage across generator and scale shifts that never shaped search. We introduce RegimeGym, a controlled four-domain testbed, and RegimeForge, a training-free discovery loop that exploits a useful asymmetry: model proposals are scarce, while executable falsification is cheap. CPUs mine an 80-regime atlas, convert diverse failures into compact behavior dossiers, and direct a global explorer plus a targeted repair. Across three model families and 12 paired runs, RegimeForge improves all nine nonzero cell effects and ties three; every domain and model family has a positive mean. Median hidden-regret reduction is 1.06 points, the library-independent geometric objective falls 6.82%, and mean regret drops from 22.43% to 12.74%. Frozen programs evaluated on 320 newly drawn hidden regimes retain 9/2/1 positive/tied/negative cells, a 1.84-point median, a 8.85% geometric gain, and a positive crossed-axis interval. Crossed mechanism replications show a 46.60-point full-atlas effect in GLM/cache and a 2.12-point, 4/4-positive effect in GLM/load. Matched removal gives behavior deltas a 4.04-point median advantage in 6/8 paths; ranked and random family-diverse dossiers improve all 16 matched arms. The 84-call main study costs $4.73 because models propose while CPUs falsify. Together, these results establish executable falsification as a practical, affordable foundation for agentic algorithm discovery.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.