Training-Free and Generation-Free Answerability Estimation in Foundation Models
Abstract
Foundation models, including large language models (LLMs) and vision-language models (VLMs), answer some queries correctly and fail on others. We study this through the concept of answerability: whether a model can correctly answer a given query. Estimating answerability before generation enables adaptive inference, such as model routing, retrieval, or abstention. Most existing methods either estimate answerability after generation or rely on task-specific data to construct an estimator. We consider a training-free and generation-free setting, in which the estimate requires neither parameter optimisation nor answer generation. We propose Answerability Direction Probing (ADP), guided by a decomposition of answerability into input adequacy, knowledge boundary, and reasoning capability. ADP constructs answerability-related concept directions from vocabulary tokens in the unembedding space and probes representations across layers, using a Jacobian lens for intermediate-layer alignment. The score is obtained from a single forward pass. With a single configuration shared across LLMs and VLMs, ADP outperforms the compared training-free and generation-free methods across multiple model families and diverse tasks. In a model-cascading case study, an ADP-based cascade matches the accuracy of always using the larger model while invoking it less often.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.