acceptodds
Under review as a conference paper at ICLR 2027

Same Risks, Different Error Bars: Disagreement-Aware Uncertainty for Cross-Validated Model Selection

Abstract

A practitioner who chooses between two fixed classifiers by cross-validation reports the selected model's cross-validated score, usually with a fold-based or binomial error bar. Fold complements overlap heavily, so the selection is re-run on almost the same data in every fold, and when the two models are close the folds flip between them. We show that this adds a reuse term to the mean squared error of the score that the usual error bars omit, and that the term is set by how the two models make mistakes jointly, not by their risks alone: pairs with identical risks can differ in it fifteen-fold. A pathwise identity (the cross-validated score equals the winner's empirical risk plus 1/n times the margins of the folds whose selection flips) yields an exact finite-sample evaluator of this error for any law of the paired 2x2 loss table and any fold count dividing n. Applied to the observed table it gives a reuse-aware root mean squared error (RMSE), which apart from a floor on empty cells is the ideal paired-loss bootstrap of the select-and-score pipeline for the two fixed models, computed without resampling; its coverage is measured, not guaranteed. On a synthetic grid fixed before any run the usual intervals cover as little as 75.0% at nominal 95%, while the reuse-aware interval's worst cell is 0.895 against 0.857 for a constant widening matched to its width. On 252 cells built from open-weight language-model checkpoints its worst cell is 0.907, which an exactly calibrated interval would fall to or below with probability 0.07 (expected worst cell 0.917), against 0.877 for the binomial and Agresti-Coull intervals (probability below 1e-5). On our declared grid of 1212 cells of real fitted classifiers, however, a textbook Agresti-Coull interval does as well or better, mostly at n = 200 where the estimated table is noisiest; at n = 1000 the reuse-aware interval's worst cell is the higher, 0.898 against 0.878. In practice, with two fixed candidates, report the winner's empirical risk, which the identity recovers from a cross-validated score and its fold contrasts; a cross-validated score that is reported should carry the reuse-aware RMSE, not a fold-based or binomial error bar.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.