Improving Fuzzy Robustness in LLM Safety with Automated Red Teaming
Abstract
AI safety and alignment must robustly stay ahead of capabilities to mitigate significant risks. In this paper, we propose a scalable approach to generate adversarial safety data and show that reinforcement learning (RL) training improves the robustness of large language models (LLMs) on safety in biology and cybersecurity. Motivated by robust optimization with computationally bounded adversaries, we scale the test-time compute of reasoning models through a feedback loop harness that allows them to iteratively propose attacks, query target LLMs, and refine candidates using safety judge feedback. We show that reasoning models without specialized red-teaming training can already achieve high attack success rates (ASR) against frontier LLMs, while trained attackers perform better still. To mitigate reward hacking from fuzzy safety feedback, we use the worst@n attacker objective, which requires success across repeated evaluations, and show that increasing yields jailbreaks that are more reliable and transferable to a held-out defender LLM. Finally, we apply automated red-teaming to generate a large set of adversarial tasks in cybersecurity and biology and incorporate these robustified tasks into RL post-training. In a controlled comparison, training with robustified tasks improves safety's robustness on held-out fixed jailbreaks from 45% to 68% and on adaptive online attacks from 71% to 92%. In sum, training on robustified tasks generally improves fuzzy robustness to new tasks and attacks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.