acceptodds
Under review as a conference paper at ICLR 2027

Thinking Longer, Refusing Less: Reasoning Shifts What Models Accept

Abstract

Thinking longer is expected to help reasoning models catch what is wrong with a request, but it may instead talk them into accepting it. To find out which, we cut the natural reasoning of ten open-weight models at several points, so only how much of one trajectory precedes the answer varies, and record whether each answer accepts harmful and benign requests, false and true premises, and wrong and correct proposed answers. Thinking longer shifts what models accept. They accept more harmful requests and false premises, most in the longest trajectories, while they refuse fewer benign requests and accept no more proposed wrong answers. In nearly every trajectory that ends in wrong acceptance, the model writes a reason to accept, usually an assumption of its own, and removing this reason lowers wrong acceptance. Motivated by this observation, we propose Stance-Contrastive On-Policy Self-Distillation (SC-OPSD), which distills into the weights a verification that the same model writes without seeing its reasoning, out of reach of such reasons. On Qwen3.5-4B and OLMo-3-7B-Think, SC-OPSD accepts fewer harmful requests and false premises than the base model on every source, especially where the base model thinks longest. Its harmful acceptance changes little as it thinks longer and at 80% of its thinking stays below the base model's at 20%, demonstrating that thinking longer need not mean refusing fewer harmful requests.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.