Decodable but Unresolved: Chain-of-Thought Conflicts in Reasoning Models
Abstract
A reasoning trace can support a correct answer to the wrong question. We study whether reasoning models preserve the user question when a supplied trace addresses another. Our diagnostic, CoT-Swap, first verifies that a model answers both questions correctly with matching traces, then swaps the trace and scores the output against both answers. Across factual and mathematical tasks, reasoning-tuned models frequently follow the trace's question, even when further reasoning is allowed. Yet mismatch remains linearly decodable on unseen question identities. Complementary training and state-intervention studies test what changes this behavior. Under matched training recipes, trace-token supervision increases redirection, while conflicting-prefix supervision attenuates it. Donor controls show that recovery depends more on sharing the user question than on sharing its answer alone. A low-rank map learned to predict matching-minus-mismatched state differences outperforms probe-direction steering in matched comparisons and partially restores user answers on unseen questions and new tasks, without test-time donors. The results distinguish answering ability, mismatch readout, and successful correction: neither demonstrated competence nor decodable conflict ensures that generation follows the user question.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.