Think Before You Lie: How Reasoning Leads to Honesty
Abstract
As Large Language Models' (LLMs) capabilities increase, deceptive behavior threatens to bottleneck safe and reliable deployment. We observe a surprising phenomenon: forcing models to reason before answering systematically increases their honesty in moral dilemmas. This result is consistent across three families of models and multiple sizes (for a total of 7 LLMs tested), and across two datasets released with this paper. However, both autoraters and human raters are unable to predict whether an LLM will be deceptive from the content of the reasoning traces. Moreover, deceptive behavior is also less stable than honest behavior under three diverse forms of noise: paraphrasing the input, resampling the answer, and injecting noise into the activations. We consider three potential root causes for these findings: inherent dataset bias, a mismatch of confidence between honest and deceptive choices, and a landscape explanation. Our experiments refute the first two hypotheses, but corroborate the landscape explanation: deceptive attractors are narrower and shallower attractors than honest attractors. Reasoning trajectories are then interpreted as a noisy traversal of space, nudging the answers toward the more stable honest regions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.