Resampling Robustness Does Not Establish Comparator Adequacy: A Leakage-Free Hyperspectral Benchmark
Abstract
Empirical machine-learning benchmarks quantify uncertainty over seeds, splits and sampling masks, but these estimates are conditional on a fixed benchmark specification: a preprocessing pipeline that must respect the hold-out, and a comparator set that defines the reported margin. We study both assumptions on identical stored runs from a hyperspectral reconstruction benchmark. First, we formulate hold-out invariance and show that smoothing spectra before the hold-out is chosen violates it: a 7-point Savitzky–Golay filter contaminates 99.6% of nominally observed channels under an interleaved split and 37.5% under a binned variant. In a fixed-target control, this contamination produces heterogeneous error reductions, with no median effect under a spatial hold-out the spectral smoother cannot contaminate and reductions as large as 46% in individual method-and-cube cells. Second, we define comparator-set sensitivity and measure it from stored per-run errors. An additive Gaussian field scored against two comparator baselines, linear interpolation and PCA, is best on 11 of 12 cubes by 12.5% (7.7–17.1% hierarchical bootstrap interval), yet against the full evaluated pool of 15 comparators it trails the strongest method by 24.4% (16.8–33.4%). The same reversal occurs for SIREN, whereas Soft-Impute remains sign-stable. Across comparator subsets, variation can cross the scientific decision boundary even where repeated sampling yields a narrow interval conditional on the restricted suite. We therefore propose a Comparator-Set Robustness Audit that reports how a claim depends on the evaluated comparator pool, while treating subset agreement as a descriptive sensitivity analysis rather than a probability over future benchmark choices.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.