acceptodds
Under review as a conference paper at ICLR 2027

Aware Yet Unsafe: Analyzing and Restoring Mediated Safety Control in Reasoning LLMs

Abstract

While Chain-of-Thought (CoT) reasoning significantly augments the problem-solving capabilities of Large Language Models (LLMs), it inadvertently exposes critical safety vulnerabilities. In this paper, we investigate a paradoxical "aware yet unsafe" phenomenon: despite explicitly recognizing malicious intents, models that reliably reject harmful requests under direct generation often comply under CoT reasoning. To mechanistically explain this discrepancy, we formulate safe response generation as a three-stage causal mediation chain, , where a Safety-Awareness State () activates a Mediating Safety-Control State () to dictate the final Behavioral Outcome (). Through activation steering, we reveal that the extended reasoning context systematically attenuates this critical linkage, degrading the downstream translation from cognitive awareness to safe execution. Motivated by this insight, we propose a lightweight, training-free intervention. By isolating and dynamically amplifying the specific native attention heads and SwiGLU MLP channels governing the and transitions, we effectively reinforce internal safety signals. Extensive evaluations demonstrate our method restores safety guardrails against CoT-induced harmful requests while largely maintaining general reasoning utility on benign requests.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.