acceptodds
Under review as a conference paper at ICLR 2027

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection

Abstract

Most existing LLM safety evaluation and defense methods follow a static formulation: jailbreak vulnerabilities are evaluated with fixed attack methods, and guardrails are trained on fixed malicious prompt datasets. However, real-world adversaries continuously evolve their attack capabilities and expand the attack space. To address this challenge, we propose DARWIN, an evolutionary attack-defense framework that formulates jailbreaking as a continual evolutionary process and continuously updates guardrails through an attack-defense loop. We propose DARWIN-Attack as an evolutionary adversary that expands its attack capabilities through strategy discovery, mutation, and selection. DARWIN-Attack discovers new attack strategies from broad external sources, generates new variants through self-reflection and genetic evolution, and filters effective strategies according to their performance against aligned LLMs. During the attack execution phase, DARWIN-Attack adaptively selects and composes evolved strategies according to feedback from target LLMs and guardrails. Through continuous evolution, DARWIN-Attack achieves state-of-the-art attack success rates against frontier LLMs and guardrails, e.g., nearly 100% on DeepSeek-V4-Pro, over 90% on GPT-5.5 and nearly 100% on YuFeng-XGuard, outperforming recent evolving frameworks such as LSA and MAGIC. The continued evolution of DARWIN-Attack exposes new safety vulnerabilities and requires timely corresponding updates to safety defenses. Therefore, on the defense side, we introduce DARWIN-Guard, an online adversarial guardrail training paradigm, which iteratively trains the guardrail on the emerging adversarial samples generated by DARWIN-Attack. To improve robustness without sacrificing utility, DARWIN-Guard jointly learns from malicious and benign disguised queries, encouraging the guardrail to recognize underlying intent rather than superficial attack patterns. Through continuous evolution, DARWIN-Guard achieves an average unsafe recall of 95.0% across multiple safety evaluation benchmarks, outperforming recent advanced guardrails such as YuFeng and Nemotron. Meanwhile, DARWIN-Guard maintains an average benign pass rate of nearly 100% on standard benign benchmarks and achieves the best performance on over-refusal benchmarks. Our code and model are available at https://anonymous.4open.science/r/DARWIN-D8A6.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.