The Other Side of Answering: How Language Models Respond Beyond Direct Answers
Abstract
Large language model (LLM) abstention studies when models should answer a question and how they should respond when a direct answer is unwarranted. Existing work typically evaluates whether models follow prescribed actions for predefined problem types, leaving no unified account of why direct answering is unwarranted or how alternative valid responses should be evaluated. We instead evaluate whether a response exceeds what the question supports. We formulate four ordered necessary conditions for direct answerability: well-definedness, knowledge support, criterion invariance, and system capability. Each failed condition specifies what a response must not assume or claim. Guided by this framework, we construct a diagnostic benchmark of 3,474 manually screened questions and evaluate seven models. Our results reveal substantial differences across failure layers, with violations most frequent for underspecified questions and inaccessible information. Boundary adherence and helpfulness emerge as related but distinct abilities. Ordinary CoT has inconsistent effects, whereas an explicit four-condition check reduces violations across all layers, especially from 26.5% to 11.8% for well-definedness and from 29.7% to 11.3% for system capability, without reducing helpfulness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.