Defect-Aware Confidence Routing: Black-Box Confidence Signals Have Regime-Dependent Reliability
Abstract
Reliable LLM deployment requires knowing when to abstain. In the black-box setting, practitioners combine confidence signals—verbalized confidence, self-consistency, cross-model dissent—using a single fixed rule. We show this is fundamentally miscalibrated: the reliability of each signal is regime-dependent. Cross-model dissent is an excellent error indicator on unanswerable questions (AUROC up to ) but collapses to near-chance () on answerable ones, where it no longer separates errors, so any fixed fusion is trapped in a cross-regime compromise—and indeed degrades below the best single signal. We turn this into a method, Defect-Aware Confidence Routing (DACR). Two facts make it deployable: the polarity depends on the latent defect type, and the defect type is predictable from the question text by a light probe (AUROC ) that leaks no labels. DACR uses this input-side probe to route signals—selecting which signal and with which polarity—through probe signal interactions. End-to-end, DACR significantly beats the best single signal by points in-domain and points under zero-transfer to SelfAware (paired bootstrap CIs exclude zero), recovering of the per-slice oracle gap; the same gate is diagnostic, predicting on which corpora the router transfers directly and on which its defect structure differs enough to require re-fitting. A five-seed ablation isolates the gain to routing (the same features without routing interactions degrade below behaviour-only), and leave-one-model-out confirms it generalizes to five of six held-out models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.