Beyond Mechanisms: An Event-Level Characterization of LLM Jailbreaks
Abstract
Recent work has uncovered multiple internal mechanisms underlying jailbreaks, including changes in harmfulness recognition, refusal, and so on. Mechanism-level perspectives, however, are inherently fragmented: different studies typically isolate different factors or processes involved in jailbreaks, often at different levels of abstraction. As a result, they provide important local explanations of jailbreak behavior, but do not by themselves characterize how these factors jointly change when a jailbreak succeeds. We instead shift the unit of analysis from mechanism to event, asking which underlying factors, taken together, characterize the successful jailbreak event. Because reduced refusal does not necessarily imply substantive assistance toward the harmful goal, we treat harmful compliance as a separate outcome factor. To this end, we first investigate how harmfulness recognition, refusal, and harmful compliance jointly change during a successful jailbreak event, and find a consistent pattern across three LLM families: harmfulness recognition remains largely unchanged, while refusal decreases and harmful compliance increases. This event-level finding naturally motivates us to formulate jailbreak suppression as a constrained problem: increasing refusal and reducing harmful compliance under the constraint that harmfulness recognition remains approximately unchanged. We then implement this formulation with a constrained restoration method. Experiments across three LLM families show that the resulting method consistently reduces jailbreak success.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.