Pretending to Think: Unmasking and Eliminating Post-Hoc Rationalization in LLM Safety Chain of Thought
Abstract
Chain-of-thought safety alignment promises reasoning-driven guardrails, but a readable analysis does not by itself show that evidence determines the final de- cision. We separate early decision predictability from evidence-responsive updat- ing and evaluate both on matched safety and general-task prompts. The resulting Safety-as-Reasoning (SaR) objective combines an evidence-conditioned symmet- ric confidence constraint with evidence-first paired supervision. In the reported evaluation, the revised objective reduces harmful unsafe compliance to 4.3% and benign over-refusal to 9.4%, while retaining 82.0% general-task accuracy. Cor- rect refusal-to-compliance and compliance-to-refusal updates reach 73.8% and 77.0%, respectively. The framework treats early tendencies as useful hypotheses and trains later decisions to respond to validated evidence in either direction
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.