acceptodds
Under review as a conference paper at ICLR 2027

Error Bars Change the Leaderboard: Handcrafted Features Are Not Beaten in Automated Alpha Mining

Abstract

A benchmark can order methods it has no power to separate, and the ordering can depend on choices its table never reports. We measure how far both go on automated alpha mining, the search for formulas that predict the cross-section of stock returns, which has its own machine learning leaderboard. Miners built on genetic programming, reinforcement learning, GFlowNets and LLM agents are ranked by mean information coefficient (IC) on a fixed test period. Some report dispersion across search seeds, none reports uncertainty over time, and none tests whether the ordering between methods is significant. We re-run the released code of AlphaGen, AlphaSAGE and a genetic-programming baseline, at their published budgets and at a matched one, together with AlphaAgent's released formulas, the Alpha158 handcrafted library and the Kronos foundation model zero-shot, under one point-in-time protocol on CSI300 constituents. No miner is ever the unique survivor of a ninety percent model confidence set, at either horizon or on any window. At twenty days, the horizon they are trained for, a six-factor style model has the highest mean IC, 0.0768 against 0.0716 for its closest competitor, and every searcher run at its released budget falls outside the set. At one day, the label these benchmarks score on, the set retains a single handcrafted model, which kills one of our own pre-registered claims and we report the kill. The ordering it publishes then fails three separate checks. Only 16 of 66 pairs separate under familywise control, and among methods that sit next to each other in the ranking, 1 of 11 do. Switching to an information ratio or a cost-adjusted Sharpe drops the style model below the handcrafted library and below a mined pool, because its IC climbs with the forecast horizon while fitted signals peak and decay. Spending AlphaGen's released budget instead of a matched one buys +0.0047 training IC and loses -0.0080 out of sample. We also pre-registered an online false discovery layer over candidate streams; it works in simulation, abstains on real streams, and fails its own test, which we report as a negative result. Pools, candidate streams, pre-registrations and a 19-check verification suite are released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.