Hidden Reasoning Changes KV States More Than It Changes Behavior
Abstract
When a reasoning language model deletes its hidden trace before the next conversational turn, the answer's cached key–value states become stale: they were computed at the wrong positions and under the influence of text that no longer exists. Serving systems handle this by recomputing the answer's cache from the visible transcript, at a cost linear in answer length on every follow-up. We find that most of this recomputation is unnecessary. An exact RoPE rotation fixes the positional error, and the remaining content contamination—though large in norm and more behaviorally sensitive than a random perturbation of the same size—rarely changes the model's next-token decisions, because the induced logit shifts almost never cross the margin between the top two candidates. Representational, distributional, and decision fidelity turn out to be separable: the size of the cache error does not predict which turns drift (Spearman ), but the decision margin does, and since it is readable from the reuse forward alone it serves as an online signal for selective recomputation (AUC ). The finding replicates across four models spanning dense, mixture-of-experts, and hybrid attention (– next-turn top-1 agreement), and a matched control against visible-span deletion and prompt swapping shows it is a general property of position-portable cache reuse, with hidden-reasoning deletion the hardest case. On two-turn tasks that force the model to read its own stored answer, zero-repair reuse matches full recomputation to within statistical error. For reasoning-model serving, patched vLLM experiments show median next-turn time-to-first-token speedups of and at concurrent batch sizes and , respectively, excluding prior-turn compaction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.