JailbreakProfiler: Fine-Grained Attack Pattern Recognition and Open-Set Attribution for LLM Safety
Abstract
Large language models (LLMs) face a diverse and evolving landscape of jailbreak attacks, yet most deployed defenses return coarse harmful/benign or generic policy-category verdicts rather than fine-grained attack attribution. We formulate this as open-set attack attribution: a system should assign a known fine-grained attack label when possible, but preserve uncertainty and route unresolved cases to analyst review. We present JailbreakProfiler, a framework comprising (i) an attack-pattern benchmark covering 19 fine-grained categories in seven coarse families, with 6,077 in-distribution samples and 900 unseen-wrapper samples, (ii) a two-stage profiler that performs coarse routing followed by family-constrained fine-grained recognition, and (iii) a review-queue mechanism that escalates unrecognized patterns to analysts. Experiments show that eight general-purpose LLMs are unreliable profilers: most force 64–86% of out-of-distribution attacks into known labels, silently masking novel threats. In contrast, JailbreakProfiler achieves 96.21% closed-set accuracy, routes all 450 attacks in the uncapped OOD pool to review, adds only 13.63 ms p50 offline profiling latency per request, and generalizes to external jailbreak data. In simulated deployment over drifting traffic, it tracks per-window attack rates with Pearson r ≥ 0.95 and flags novel wrappers in the window they first appear. Unlike prior work that treats jailbreaks as a static closed-set problem, our framework is designed for the operational reality of LLM safety, where attack patterns are unknown, evolving, and must be recognized at scale.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.