acceptodds
Under review as a conference paper at ICLR 2027

Beyond Mechanisms: An Event-Level Characterization of LLM Jailbreaks

Abstract

Recent work has uncovered multiple internal mechanisms underlying jailbreaks, including changes in harmfulness recognition, refusal, and so on. Mechanism-level perspectives, however, are inherently fragmented: different studies typically isolate different factors or processes involved in jailbreaks, often at different levels of abstraction. As a result, they provide important local explanations of jailbreak behavior, but do not by themselves characterize how these factors jointly change when a jailbreak succeeds. We instead shift the unit of analysis from mechanism to event, asking which underlying factors, taken together, characterize the successful jailbreak event. Because reduced refusal does not necessarily imply substantive assistance toward the harmful goal, we treat harmful compliance as a separate outcome factor. To this end, we first investigate how harmfulness recognition, refusal, and harmful compliance jointly change during a successful jailbreak event, and find a consistent pattern across three LLM families: harmfulness recognition remains largely unchanged, while refusal decreases and harmful compliance increases. This event-level finding naturally motivates us to formulate jailbreak suppression as a constrained problem: increasing refusal and reducing harmful compliance under the constraint that harmfulness recognition remains approximately unchanged. We then implement this formulation with a constrained restoration method. Experiments across three LLM families show that the resulting method consistently reduces jailbreak success.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.