acceptodds
Under review as a conference paper at ICLR 2027

Repeat or Retrieve? Local Copies Suppress Remote-Source Sensitivity

Abstract

Reasoning traces often state an intermediate value twice: once where it is computed and again where it is used. Does the next prediction depend on the earlier source or the nearby copy? We separate these occurrences with source-span activation patching and controlled edits to the copy. Across Gemma-2, OLMo-2, and Qwen3, a correct local copy reduces the mean absolute source-patch effect by more than % on controlled arithmetic traces. Randomized trace formats and length-matched filler controls show that the reduction depends on copied values. Controlled training shows that this preference depends on the reliability of copies. In small transformers trained from scratch, corrupting one operand on half of the restated training lines reverses the preference at fixed restatement frequency: copy sensitivity becomes negligible while source sensitivity remains. When copies are reliable, making them more frequent also strengthens source suppression. The preference affects generation: a conflicting copy supplied in an arithmetic prefix makes Gemma-2 and Qwen3 produce its implied result about % of the time, and a source change reaches a copy that the model writes itself in -% of samples. These results show how local restatements change which occurrence affects a prediction and identify training conditions that can reverse that change.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.