acceptodds
Under review as a conference paper at ICLR 2027

Context or Knowledge: How LLMs Attribute Their Own Answers

Abstract

Language models can answer a question from a supplied document or from their parametric knowledge. When later asked where the answer came from, what determines their source report? We study this by intervening on the model's key–value (KV) cache after the answer is generated and before the source question is asked. On a controlled factual task with four models (7B–32B), changing or removing the document's cache never changes the report. In contrast, transplanting the question-and-answer (Q+A) cache between runs with identical question and answer tokens transfers it: every eligible report follows the donor, while the recipient's document is unchanged. On the two smaller models, steering Q+A values along directions fitted before any source question reverses 98.9–100% of reports across four document conditions at a fixed strength, while norm-matched random directions flip at most 1.1%. The components of keys and values along fitted directions carry the transplant effect: copying them alone nearly reproduces the report change; holding them fixed nearly prevents it. Applied during answer generation, the same value directions also change which source supplies the answer, though more weakly and depending on the prompt. Source reporting in this setting thus depends causally on cached state produced while answering. We do not establish general reporting accuracy or a complete interpretation of this state.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.