Correct Trace Conclusions, Wrong Answers: Unfaithful Capitulation Under Pressure
Abstract
Reasoning models are trusted in part because they show their work. We find that the reasoning trace and the final answer can nevertheless diverge. When a user pushes back, Qwen3 models sometimes reason their way to the correct option and then deliver a different one, a failure we call **unfaithful capitulation (UC)**. Using gold-blind extraction of each channel’s final commitment, we find that 30% of first capitulations across seven multi-turn runs follow a reasoning trace that still concludes correctly. The reasoning itself is not what breaks: holding each saved trace fixed and regenerating only the answer, replacing the challenge with a neutral request, or removing the dialogue it refers to restores 93–99% correct answers. The trace helps, but neutralization works almost as well without it, suggesting that the failure lies in how the final answer responds to the conversational context. Finally, because corrupting the trace removes UC as surely as fixing the answer does, agreement between reasoning and answer is the wrong repair target: trace correctness and answer correctness must be evaluated separately.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.