Strong yet Controllable LLM Abstention without Learning to Abstain
Abstract
Large language models (LLMs) can reduce hallucination by abstaining from questions they would likely get wrong. Abstention can be implemented outside the model (selective answering) using a confidence threshold that controls the answer rate (coverage), or learned inside it (abstention training), which can perform better but fixes coverage once trained. To compare the two, we match their coverage and decompose risk (error rate) into answering ability and distinguishing ability, *i.e.*, how well the model answers questions and how accurately it abstains. The decomposition shows that abstention training performs better mainly by answering better: with strong confidence estimators, the two have nearly equal distinguishing ability. This motivates **T**rain for **A**ccuracy, then **P**ost-hoc **S**elect (**TAPS**), a modular approach that keeps coverage controllable, needs only standard post-training and an off-the-shelf confidence estimator, and benefits from advances in either component. Across three LLMs and three factual QA benchmarks, TAPS with a probing-based confidence estimator attains 1.5–9.4 pp lower mean risk than three representative abstention training methods at matched coverage. Within TAPS, accuracy training often weakens the model's own confidence signals, yet refitting the estimator largely preserves distinguishing ability. Together, these results establish TAPS as a strong yet controllable baseline and raise the bar for future abstention training: sacrificing control over coverage should yield gains beyond what can be achieved by optimizing answering and abstaining separately. The source code will be made publicly available upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.