acceptodds
Under review as a conference paper at ICLR 2027

SelfShield: Unlocking LLMs' Self-Protective Capability Against Jailbreak Attacks via a Resilient Anchor Layer

Abstract

Large language models (LLMs) have achieved remarkable performance across diverse tasks, yet remain highly vulnerable to jailbreak attacks. Existing work has primarily focused on filter-based, steering-based, and decoding-based methods to defend against jailbreak attacks at inference time. However, these approaches largely overlook the intrinsic robustness of LLMs: their hidden representations may retain self-protective capability even when the model is successfully jailbroken. In this work, we systematically investigate how hidden states evolve across layers under jailbreak attacks. A key observation is that, although jailbreak attacks substantially disrupt the model's internal safety discrimination across most layers, a resilient intermediate layer can still preserve strong separability between jailbreak and benign inputs. Motivated by this observation, we introduce SelfShield, an inference-time defense that exploits this safety-resilient layer as an internal anchor to guide subsequent generation. First, SelfShield employs a carefully designed probe network to track layer-wise harmfulness changes and locate the safety-resilient anchor layer. Second, SelfShield adaptively injects steering vectors only into the layers following the anchor layer to counteract the downstream degradation of safety discrimination. Finally, SelfShield uses the anchor-layer token distribution to recalibrate the output token distribution, allowing the model's preserved intrinsic safety capability to guide safer generation. Experiments across multiple jailbreak attacks and utility benchmarks demonstrate the effectiveness of SelfShield, which significantly improves the safety of LLMs while preserving their general capabilities. Code is available at https://shorturl.at/BuEPL.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.