FORGE: Evidence-Governed Algorithm Discovery by Attributing Failures and Certifying Interventions
Abstract
Algorithm-discovery agents traditionally iterate through execution, scoring, and revision; however, the research state guiding each revision remains either an unstructured narrative or a unidimensional ranking—a practice that systematically obscures failure origins, renders interventions non-falsifiable, and actively incentivizes evaluator overfitting. We introduce FORGE, an evidence-governed agent that maintains an append-only state comprising executable outcomes, counterexamples, failure taxonomies, resource consumption, and full provenance. At each iteration, FORGE isolates a specific failure mechanism, articulates a predicted behavioral shift, compiles a bounded candidate edit, subjects it to robustness testing without access to hidden labels, and accepts modifications only after deterministic validity and resource constraints are satisfied. Across a 640-cell Exp2 benchmark, FORGE attains the highest mean hidden score (0.910), outperforming AIScientist v1 (0.892), v2 (0.635), and OpenEvolve (0.765), with a +0.017 gain [95% CI: +0.004, +0.030] over the strongest baseline. In a comparator-network repair study, FORGE yields 26 verified repairs out of 40, compared to only 5 for baselines. Two exploratory transfer tasks further confirm generalization beyond synthetic search: in memory retrieval, FORGE achieves 0.742/0.751 vs. FullText (0.604/0.676) and NaiveRAG (0.650/0.719); in skill optimization, FORGE r55 attains 0.867/0.811 on SearchQA/SpreadsheetBench, surpassing SkillOpt (0.831/0.625) by +0.036/ + 0.186. These results position FORGE not merely as a stronger optimizer, but as a principled framework toward interpretable, auditable, and increasingly autonomous algorithmic discovery—with clear potential to scale across domains where reliable evidence, rather than heuristic ranking, becomes the cornerstone of intelligent search.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.