acceptodds
Under review as a conference paper at ICLR 2027

Out of Sight, Still in Mind? Separating Internal and Answer-Level Effects of Retained Key-Value History

Abstract

Cache eviction, prompt compression, and context management change which source tokens remain directly accessible in the key–value (K/V) cache, while later states computed from those tokens can remain. How removed text then influences computation, and whether internal effects translate into answer-level effects, is a measurement question. We study it with a paired cache-interchange design. Two histories specify the same task, by worked examples or by a direct instruction, and are followed by an identical bridge text; we replace the source-prefix cache with a shared neutral cache and swap the retained bridge states between histories under an identical query. We read donor projections at three levels, an intermediate hidden state, the next-token logits, and the correct-minus-incorrect answer margin, with endpoint-specific aggregation, and we report candidate choices wherever candidate scores were archived. Across five arms (Qwen, Mistral, and OLMo; two synthetic rule tasks; 500 indirect-object-identification (IOI) episodes), retained bridge states remain causally active after source-prefix neutralization: mean net donor-axis projections at the intermediate readouts range from 0.59 to 0.97, all 1,780 episodes are positive, and the format difference vanishes when bridge attention is cut. These internal effects do not establish corresponding answer-level effects. Logit-contrast projections are smaller (0.007–0.26), and answer-margin projections range from to . On IOI, which the model solves without the histories, a hidden-state projection near one coexists with a small negative donor-oriented margin shift ( [, ]) and no changed candidate choice in either patch direction. We formalize what donor projections can and cannot establish and give reporting practices that keep internal alignment, answer margins, and candidate choices separate.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.