acceptodds
Under review as a conference paper at ICLR 2027

Faithfulness Anchoring as Defense Against CoT Hijacking: A Mechanism-First Approach via Residual-Stream Analysis

Abstract

Safety-aligned reasoning models are vulnerable to Chain-of-Thought (CoT) hi- jacking, a jailbreak that prepends long benign reasoning traces before a harmful instruction, causing the model to comply despite its safety training. Understanding why this attack succeeds is a prerequisite for building principled defences, yet the internal mechanics of CoT hijacking remain unexplored. We investigate by tracking the model’s refusal signal, measured as a projection onto a refusal direction in the residual stream, across all token positions during hijacked generations. Our key finding is that benign padding does not gradually erode this signal; instead, it suppresses it uniformly while the model processes the input tokens. Attack out-come is determined by whether the signal recovers at generation onset or remains suppressed. Across three model families, we confirm that early-generation signal strength correlates with attack outcome (Pearson r ≈ −0.60, p < 0.0001, n=150 per model) and causally validate it via activation patching. Building on this finding, we propose faithfulness anchoring, a prompt-level intervention that restores the refusal signal at generation onset without any access to model internals, reducing mean ASR by 22.7pp across three models while keeping benign task accuracy close to baseline. We further show that activation-level steering cuts ASR further (mean −32.0pp) but trades capability for safety in a model-dependent way, motivating the prompt-level defence. Finally, we find the defence is not adaptively robust: attacks targeting the anchor’s own instruction recover much of the original attack surface.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.