acceptodds
Under review as a conference paper at ICLR 2027

Think Before You Lie: How Reasoning Leads to Honesty

Abstract

As Large Language Models' (LLMs) capabilities increase, deceptive behavior threatens to bottleneck safe and reliable deployment. We observe a surprising phenomenon: forcing models to reason before answering systematically increases their honesty in moral dilemmas. This result is consistent across three families of models and multiple sizes (for a total of 7 LLMs tested), and across two datasets released with this paper. However, both autoraters and human raters are unable to predict whether an LLM will be deceptive from the content of the reasoning traces. Moreover, deceptive behavior is also less stable than honest behavior under three diverse forms of noise: paraphrasing the input, resampling the answer, and injecting noise into the activations. We consider three potential root causes for these findings: inherent dataset bias, a mismatch of confidence between honest and deceptive choices, and a landscape explanation. Our experiments refute the first two hypotheses, but corroborate the landscape explanation: deceptive attractors are narrower and shallower attractors than honest attractors. Reasoning trajectories are then interpreted as a noisy traversal of space, nudging the answers toward the more stable honest regions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.