When Proxy Calibration Controls the Wrong Risk: Accepted-Set Target Mismatch in Selective LLMs
Abstract
Selective LLM systems often use inexpensive proxy or judge labels for calibration, but a statistically valid certificate only controls the risk defined by those labels. Does proxy-valid selective calibration control the target risk that deployment actually cares about? We derive an accepted-set identity decomposing the proxy–target risk gap into signed disagreement contributions; global judge agreement alone does not establish transfer. At α = 0.20, the held-out main evaluation shows proxy-label calibration selecting λ = 0.2 with coverage 1.000 and held-out target risk 0.250, whereas target-label calibration selects λ = 1.0 with coverage 0.591 and target risk 0.113. A threshold-wise full-frame diagnostic shows that accepted false-positive contributions outweigh false-negative contributions, producing the proxy–target risk gap. In a 400-row blinded validation sample, human judgments agree with the reference-aware operational target on 96.75%, versus 74.50% with the reference-free proxy. The low-threshold failure reproduces under lenient rubrics across DeepSeek, Qwen, and MiniMax, whereas stricter rubrics can select more conservative thresholds or, in one case, return no solution. Increasing target-label audit budgets improves direct feasibility, whereas proxy-assisted corrections can remain conservative. Selective-risk validity is target-relative: deployment claims require measurement validation and target-specific evidence.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.