Bridging the Detection-to-Abstention Gap in Reasoning Models under Insufficient Information
Abstract
Trustworthy reasoning requires models to abstain when essential information is missing. Yet reasoning models can identify missing premises in their reasoning and still produce unsupported final answers. We study this detection-to-abstention gap by distinguishing insufficient-information detection from the subsequent decision to abstain. We propose Judge-Then-Solve (JTS), a training framework that makes answerability an explicit commitment before solution generation. The model first audits the available premises, then learns to either continue solving or end its reasoning with abstention according to its judgment. Supervised warm-up and reinforcement learning train this behavior through structured rewards and conditional length shaping. Experiments on a dense and a mixture-of-experts reasoning model show that JTS improves overall abstention and Abstention@Detection (A@D), while shortening responses, relative to a baseline that also uses supervised warm-up, reinforcement learning, and length shaping. Across three Qwen training seeds, JTS achieves overall abstention and A@D. These gains come with lower answer rates and accuracy on answered questions, while detection errors continue to limit overall abstention. Our results support explicit answerability commitment as an effective way to improve how reasoning models act on recognized insufficiency, a specific requirement for honest and trustworthy reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.