acceptodds
Under review as a conference paper at ICLR 2027

Beyond One-Shot Jailbreaking: Cross-Instance Experience Accumulation via Reusable Skills

Abstract

Existing automated jailbreak methods repeatedly rediscover similar attack structures for each new request. A successful trajectory may reveal a reusable strategy—role-playing, scenario framing, or instruction override—yet this knowledge is typically discarded once the attack terminates. This makes large-scale red teaming query-expensive and prevents experience acquired on one attack from improving the next. We ask: (RQ1) Can successful attack trajectories become reusable cross-instance memory? (RQ2) Can such memory improve attack success while amortizing search cost? We propose a reusable skill framework that addresses both questions through three mechanisms: skill abstraction (distilling successful trajectories into compact templates), multi-factor retrieval (matching skills to new prompts via lexical scoring), and library maintenance (periodic pruning and merging to prevent skill dilution). Through 217 ablation configurations and 120 transfer evaluations across 4 models and 5 datasets, we demonstrate that the skill system provides +22.7 percentage points (pp) improvement over baselines, while an optional evolution mechanism contributes an additional +0.9pp refinement. When initialized with strong prior templates, the system achieves 99.7% ASR on same-family models. Skills transfer effectively within model families (85–100% ASR across Qwen3 variants) and significantly outperform baselines on cross-family transfer (+17.5pp over AutoDAN on GPT-OSS-20B). Our results validate that cross-instance experience accumulation is both feasible and effective for systematic jailbreak prompt generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.