To Abstain or to Re-Rank? A Regime Law for Risk-Controlled Generation
Abstract
Risk-controlled generation returns a prediction set of candidate answers while controlling the probability that none is admissible. Existing methods calibrate sampling and filtering rules to satisfy this constraint, but the smallest prediction set permitted by the risk budget is not characterized. We study this efficiency problem directly. Classical selective prediction gives the optimal one-answer rule, and its instantiation here yields a simple regime law. Let be the Bayes top-1 error and the target miscoverage risk. If , a valid one-answer predictor has slack and average prediction-set size is reduced by abstaining on low-confidence questions. If , no one-answer predictor is valid and efficiency depends on ranking multiple candidates. We also show that a score carrying only within-question rank cannot selectively abstain because it lacks a confidence scale across questions. These results motivate ACF-LV, a single-stage conformal filter with a calibration-gated verifier. On three QA datasets and five language models, the high-probability single-stage variant gives smaller prediction sets on 13 of 15 settings. Under a unified end-to-end protocol, ACF-LV reduces mean APSS from 4.389 to 1.889, while all eight testable predictions of the regime law agree with the observed effect of re-ranking. Code: https://anonymous.4open.science/r/acflv-A6E8/README.md
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.