A Cross-Validation Gain Does Not Identify the Value of Its Selection Information
Abstract
Cross-validation changes two things at once. It gives the selector more validation information, and it usually changes the deployed predictor, because the fold-trained models are averaged. A gain from the complete pipeline cannot say which change produced it. We separate the two with an output-matched design that varies one component at a time, and add a third component that often travels with them: a small labelled panel that confirms the final choice. The evidence is a replay of four TabArena model families on 51 tasks, planned and sealed before any of the 765 held-back outer splits was opened. The value of selection information depends on what it is added to. For the validation-selected predictor, full out-of-fold selection raises ROC-AUC by 0.91 points when a single fold model is deployed but by only 0.28 points once fold outputs are averaged, while averaging alone is worth 0.98 points; both components lower normalized RMSE on all 13 regression tasks. An ordinary average of separately selected holdouts recovers about three quarters of the AUC gain of cross-validation. The same change of workflow matters far more for a benchmark than for deployment. RealMLP leads CatBoost on most splits of 16 tasks under a single holdout and of 32 tasks under cross-validation with averaging, yet validation picks RealMLP in 86 to 88 percent of splits either way. A confirmation panel of 32 labels that re-ranks all 800 saved models lowers AUC by 1.24 points on every binary task, and a shortlist removes most of that damage. A controlled study with actual refitting shows that the selection advantage survives full-data refitting and puts a cost on each policy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.