Semantic Support Shapes Reliability Evaluation and Downstream Decisions
Abstract
Semantic conditional reliability asks whether one predictor harms or corrects another within a specified set of true classes. Existing AUC decompositions and model-update methods do not isolate how semantic support affects reliability-scorer selection. Our support decomposition shows that predicted-label membership resolves every harmed–corrected comparison involving a query boundary, leaving only interior–interior comparisons unresolved. The resulting Semantic-Membership Baseline (SMB), normalized interior evidence, and support-order loss separate support-induced ranking from discrimination between unresolved harmed (H) and corrected (C) cases. We further characterize when semantic queries require incompatible rankings and connect this evaluation distinction to scorer selection, fallback decisions, and downstream utility. Across image classification, medical imaging, audio, intent classification, domain shift, and collaborative occupancy prediction, support-induced ranking is substantial but heterogeneous. Positive interior evidence can coexist with raw conditional AUROC either above or below SMB, and practical reliability scores can be ranked differently by raw AUROC and interior evidence. Support-aware evaluation also has an operational consequence: selecting scorers by interior evidence rather than raw conditional AUROC can change downstream fallback decisions and improve held-out utility.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.