Understanding the Knowledge–Behavior Gap in Language Models through Answer Measurement
Abstract
The knowledge-behavior gap refers to the discrepancy between the answers a language model's internal representations support and the answers it actually gives. Prior work identifies this gap by comparing what probes recover from hidden states with what models answer, and shows that probing can partly mitigate it. However, how the model's answering process loses internal answer information remains unclear. We develop a measurement theory of answer support, the total probability mass of continuations expressing the same answer. Answering estimates this support through probability readings and commits to the answer with the highest reading. Two measurement properties are key: the measurement unit, which determines whether readings concern complete answers or individual expressions, and the measurement scale, which determines whether they share a common reference. Mismatch can distort support ordering, preventing supported answers from becoming commitments. Accordingly, we construct native measurements that vary in measurement unit and scale. We further prove that, when answer support is measurable from complete-answer representations, a shared map places candidates formed along separate generation paths on a common answer-level scale. Across four frozen models, aligning either property improves answering accuracy and robustness to answer rewording; aligning both yields the highest accuracy throughout and the lowest flip rate in most settings. Supervised by the model's own support ordering, shared linear maps improve native measurements in nearly all settings and generalize even when trained only on support relations among incorrect answers. Used as generation feedback, the learned readings also improve free-form answers. Measuring answer support and converting it into commitments thus offers a route to understanding and narrowing the knowledge-behavior gap.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.