acceptodds
Under review as a conference paper at ICLR 2027

Defending Large Language Models Against Diverse Jailbreak Attacks via Multi-Concept Detection

Abstract

Despite extensive safety alignment efforts, Large Language Models (LLMs) remain highly susceptible to jailbreak attacks. Although various defenses have been proposed, their effectiveness remains limited against emerging or adaptive threats, largely because the underlying mechanisms of jailbreak attacks remain poorly understood. To elucidate these mechanisms, recent studies leverage the Linear Representation Hypothesis to extract jailbreak concepts from LLM representations. However, our analysis reveals that no single jailbreak concept generalizes across diverse jailbreak techniques, as different attack strategies exhibit distinct representation patterns. To address this limitation, we propose Multi-Concept Detection (MCD), a lightweight defense framework that robustly identifies jailbreak attempts by extracting their multi-concept features. Extensive experiments on five open-source LLMs demonstrate that MCD consistently achieves superior detection performance and remains robust against a wide range of attack strategies.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.