acceptodds
Under review as a conference paper at ICLR 2027

Feedback Loops Invalidate Significance Tests in LLM Alpha Mining

Abstract

LLM alpha mining, like other automated discovery systems, proposes candidate formulas and scores them on historical data. The scores are fed back to guide the next proposals. The system then presents its highest scoring formula as a discovery. It backs that claim with a multiple testing correction, which promises at most a 5% chance that the discovery is false. That promise assumes the candidates were chosen without looking at the data. A feedback loop breaks this assumption, because every new proposal depends on how earlier ones scored. As a result, the more these systems search, the less their discoveries can be trusted. Prior work established this risk in theory, but its remedies change how the search runs. We measure the problem in real LLM discovery loops and fix it without changing them. First, we prove that no correction computed from the chosen candidates can keep its promise for every feedback loop. For a natural class of loops, the chance of a false discovery tends to 100% as the search grows. Second, we measure the effect on stock market data shuffled so that no formula can predict returns. There, the strongest standard test for data snooping declares an LLM alpha miner's best formula significant 15% of the time. The Benjamini-Hochberg procedure does so 29% of the time. A search that never sees its scores, and the same loop with its scores shuffled, both stay near 5%. The feedback is therefore the cause. Four times the budget lifts the two rates to 28% and 70%, while both controls stay flat. The gap widens as the search grows. The same failure appears in symbolic regression, outside finance. Third, we give a fix. Counting candidates more carefully does not help. Instead, we rerun the unchanged search many times on shuffled data and set the significance threshold from those runs. This keeps the 5% promise for any searcher. On later market data the search never saw, 92% of the formulas it accepts remain significant. Its main cost is rerunning the search, and a simple statistical shortcut cuts the reruns needed from about 100 to about 20. Practitioners can therefore attach a trustworthy significance claim to any discovery loop, at modest cost and without redesigning it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.