acceptodds
Under review as a conference paper at ICLR 2027

JailbreakSkill: Scaling Automated Red-Teaming with Reusable and Ever-Evolving Skills

Abstract

Automated red-teaming has produced a growing collection of attack strategies, yet they typically remain scattered across prompts and workflows, making them difficult to systematically integrate, reuse, and improve at scale. We introduce JailbreakSkill, a skill-centric framework for scaling automated red-teaming through reusable and continuously evolving attack capabilities. JailbreakSkill packages existing attack strategies into modular, agent-ready skills that can be directly reused and adaptively selected across tasks and target models. Beyond reuse, it closes the loop between attacking and learning: attack experience is used to diagnose, refine, combine, and discover new skills, which are added back to an ever-growing skill library. This evolution lifts macro-average attack success rate (ASR) by 17.5 percentage points on AdvBench and 13.5 points on HarmBench, including a 48.6-point gain against GPT-5.4 on AdvBench, while yielding novel attack strategies such as reframing a direct request as an unfinished document-completion task. Several evolved skills also generalize to unseen prompts and target models without further adaptation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.