RT-Agent: Automated Extraction, Generation, and Adaptation of Jailbreak Attacks for LLM Guardrails
Abstract
New jailbreak techniques appear in papers, repositories, and forums faster than deployed guardrails are retrained against them. Every round costs an engineer who reads the description, reimplements the attack, builds an evaluation corpus, retrains the classifier, and checks it for regressions. We propose RT-Agent, a system that automates that path. The pipeline scrapes candidate descriptions, has a frontier LLM write an executable attack generator for each promising one and validate it against its source, scores the generated attacks against a heterogeneous ensemble of four production guardrails, fine-tunes the guard on the generated attacks, and repeats the cycle. Run over scraped documents, it promoted 30 candidates and turned seven of them into working attacks with no human in the loop, alongside six strategies reproduced from curated descriptions. The synthesized attacks raise the strict guard-jailbreak rate by up to percentage points. Fine-tuning WildGuard on seven of them in sequence lowers the share of harmful attack prompts it lets through from to on average (), suppresses each attack at its own step and partly before it, and keeps the locality on par with the base guard on standard safety benchmarks. The benign parts of the source code and data are published https://anonymous.4open.science/r/rt_agent_code-8533.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.