Router-Uniform Selective Prediction under Candidate-Set Shift
Abstract
Vision-language systems often choose candidate labels after observing an image. Even when each candidate mechanism is individually reliable, input-dependent routing can concentrate its errors and invalidate selective-risk guarantees. We characterize this failure using labeled replay: one image is evaluated under every registered mechanism. Two joint events, a selectable accepted error and an unavoidable correct acceptance, determine a router-uniform risk envelope that is exact for deterministic labels. We prove that synchronizing acceptance gates preserves worst-router coverage and cannot increase worst-router risk. The resulting objective is to predict and certify whether any mechanism errs. For set-independent scores, two nested endpoint sets additionally certify all intermediate sets. These results yield \method, which combines a learned shared gate with exact image-level binomial or remaining-pool hypergeometric bounds. On Flowers-102, a label-free router raises the risk of per-mechanism MSP gates from to . Joint replay instead certifies coverage for a DeepSets gate at risk target , while a marginal union-bound certificate for the same score rejects all inputs. A logistic gate attains mean coverage across five training seeds and improves on ViLU with the same replay certificate by – percentage points across six dataset–model systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.