acceptodds
Under review as a conference paper at ICLR 2027

An Unchanged Answer Can Hide an Operand's Causal Role

Abstract

Copying a hidden state between questions can test what information a language model uses. The final answer can conceal a contribution when another part of the computation determines which answer wins. We study this problem in Qwen3-8B on weekday and month arithmetic. We replace states at two positions using independently chosen source questions. Holding one source fixed lets us test whether the other supplies an offset or a complete result. Sources with the same offset favor the corresponding candidate even when their own results differ. This offset contribution persists in cases where every substitution produces the same answer. We trace the same candidate comparison through the receiving computation. The next attention output carries the contribution to the final position. Its effect then depends on the receiving state. Fixing the outputs of the next two feed-forward layers reduces this dependence by roughly two-thirds to three-quarters. These experiments identify an operand contribution hidden by answer recovery and distinguish its transmission from the later processing that determines its effect.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.