Stable Accuracy, Unstable Allocations:Retraining Reproducibility in Selective Prediction
Abstract
Retraining a selective predictor can change which cases receive review even when accuracy changes little. We study the reproducibility of these review assignments under exact fixed-capacity top- selection, measuring allocation reproducibility (AR) by mean pairwise normalized review-set overlap. Across six visual settings with ten retrainings each, accuracy standard deviations are 0.11–0.50 percentage points, yet 37.6–50.1% of review slots are reassigned to different cases at a 10% budget. Civil Comments/DistilBERT shows lower turnover, at 11.3%. Changes affect persistent errors and also occur when predicted classes remain unchanged. We derive a sharp bound linking Spearman correlation to top- replacements, specialize an established subset-stability identity to connect set-level and individual disagreement, and characterize how realized score and cutoff changes constrain membership changes. Two-model probability averaging reduces median allocation turnover across all tested settings and budgets, but does not eliminate it. When review identity matters, accuracy and selective risk alone do not characterize the reproducibility of the resulting allocation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.