How Corrupted Chain-of-Thought Exemplar Answers Shift Internal Answer Preferences
Abstract
A correct generated answer does not reveal whether a model ignored an incorrect exemplar answer or recovered from its influence. This raises three questions: (i) does reasoning remove the internal effect of an incorrect exemplar answer, (ii) which components carry that effect, and (iii) does that effect track final-answer accuracy? We study multi-hop comparison questions from 2WikiMultiHopQA, replacing only each retrieved exemplar's answer line with either a plausible but incorrect competing answer from its own source question or an unrelated same-category answer, and adding a reasoning–answer factorial that separates the answer-line effect from exemplar reasoning that contradicts it. We measure the target question's gold-versus-foil logit margin before and after reasoning in Llama-3.1-8B-Instruct, with a same-item cue-probe replication in Qwen2.5-7B-Instruct. Reasoning–answer inconsistency is the largest resolved effect in both models, though weaker in Qwen, and in Llama reasoning preserves it mainly through trace-following. Causal interventions identify five layer-15 writer heads and distributed upper-layer attention as shared endpoints for inconsistency and the smaller answer-plausibility effect; a low-dimensional writer signal reaches the layer-16 MLP, and query direction mediates part of the aligned answer-plausibility effect. Neither the corruption-induced margin drop nor its trace-following component detectably predicts factorial generation failure, and two manipulations differing fourfold in margin effect produce similar accuracy reductions. These results separate internal answer competition from successful output recovery.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.