acceptodds
Under review as a conference paper at ICLR 2027

Correct Trace Conclusions, Wrong Answers: Unfaithful Capitulation Under Pressure

Abstract

Reasoning models are trusted in part because they show their work. We find that the reasoning trace and the final answer can nevertheless diverge. When a user pushes back, Qwen3 models sometimes reason their way to the correct option and then deliver a different one, a failure we call **unfaithful capitulation (UC)**. Using gold-blind extraction of each channel’s final commitment, we find that 30% of first capitulations across seven multi-turn runs follow a reasoning trace that still concludes correctly. The reasoning itself is not what breaks: holding each saved trace fixed and regenerating only the answer, replacing the challenge with a neutral request, or removing the dialogue it refers to restores 93–99% correct answers. The trace helps, but neutralization works almost as well without it, suggesting that the failure lies in how the final answer responds to the conversational context. Finally, because corrupting the trace removes UC as surely as fixing the answer does, agreement between reasoning and answer is the wrong repair target: trace correctness and answer correctness must be evaluated separately.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.