acceptodds
Under review as a conference paper at ICLR 2027

Pretending to Think: Unmasking and Eliminating Post-Hoc Rationalization in LLM Safety Chain of Thought

Abstract

Chain-of-thought safety alignment promises reasoning-driven guardrails, but a readable analysis does not by itself show that evidence determines the final de- cision. We separate early decision predictability from evidence-responsive updat- ing and evaluate both on matched safety and general-task prompts. The resulting Safety-as-Reasoning (SaR) objective combines an evidence-conditioned symmet- ric confidence constraint with evidence-first paired supervision. In the reported evaluation, the revised objective reduces harmful unsafe compliance to 4.3% and benign over-refusal to 9.4%, while retaining 82.0% general-task accuracy. Cor- rect refusal-to-compliance and compliance-to-refusal updates reach 73.8% and 77.0%, respectively. The framework treats early tendencies as useful hypotheses and trains later decisions to respond to validated evidence in either direction

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.