Which OOD Detector Should You Deploy When Test-Set Gains Disagree?
Abstract
A deployed out-of-distribution (OOD) detector meets unfamiliar inputs whose mix is unknown and changing. Its effectiveness is highly sensitive to this mix. Yet on benchmarks, detectors that gain on some test sets perform poorly on others, sometimes below a vanilla cross-entropy baseline. How should one choose a detector whose performance holds across the mixes it may meet? We find that these losses are reproducible effects of training. Predictions fixed before training hold across seeds, scoring functions, architectures and datasets, with and without auxiliary OOD data. We compare a new detector with every existing choice, including the baseline under each scoring function, and compute in closed form how far the mix can shift before it stops being the top performer. At the published means of four benchmarks, the baseline wins under this comparison for about half of the training gains measured under a shared scoring function. None of those gains holds under every mix. A detector that passed this comparison held an 8.8-point lead in area under the ROC curve on new images and halved false acceptances. Gains that failed it fell short on new runs and images. OOD detectors should therefore be chosen by their stability over plausible mixes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.