Context or Knowledge: How LLMs Attribute Their Own Answers
Abstract
Language models can answer a question from a supplied document or from their parametric knowledge. When later asked where the answer came from, what determines their source report? We study this by intervening on the model's key–value (KV) cache after the answer is generated and before the source question is asked. On a controlled factual task with four models (7B–32B), changing or removing the document's cache never changes the report. In contrast, transplanting the question-and-answer (Q+A) cache between runs with identical question and answer tokens transfers it: every eligible report follows the donor, while the recipient's document is unchanged. On the two smaller models, steering Q+A values along directions fitted before any source question reverses 98.9–100% of reports across four document conditions at a fixed strength, while norm-matched random directions flip at most 1.1%. The components of keys and values along fitted directions carry the transplant effect: copying them alone nearly reproduces the report change; holding them fixed nearly prevents it. Applied during answer generation, the same value directions also change which source supplies the answer, though more weakly and depending on the prompt. Source reporting in this setting thus depends causally on cached state produced while answering. We do not establish general reporting accuracy or a complete interpretation of this state.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.