RedLoop: Maintaining Reusable Attack Strategies for Multi-Round LLM Safety Training
Abstract
Multi-round safety training must reduce harmful answers without encouraging refusal of benign queries. As training changes the defender, the attack strategies a red-teamer has stored can stop working, and its library can miss strategies that would expose the updated defender; retraining the red-teamer or rechecking every stored strategy is costly. We present RedLoop, a multi-round safety-training framework that keeps the attack-generation model frozen and adapts the red-teamer, RedScout, through a persistent library of reusable attack strategies, natural-language instructions for generating adversarial prompts. To keep the library useful without rechecking every strategy, *fitness propagation* estimates how well previous-round strategies still work from a sample of the current round's attacks and reuses the highest-ranked ones in the next training batch. Alongside strategies distilled from successful attacks, *novelty injection* contributes strategies that are distinct from the library and succeed against the defender, and later attack steps draw on them more heavily. The defender is trained with GRPO under a reward that penalizes refusal of benign queries and of benign queries carrying adversarial framing. On Nemotron-Nano-4B, RedLoop reduces HarmBench attack success from % to and XSTest over-refusal from % to %, while IFEval and MMLU stay within one point of the base model; removing fitness propagation raises HarmBench attack success to % and removing novelty injection raises XSTest over-refusal to %. Safety gains also extend to Llama-3.1-8B and OLMO-3-7B defenders.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.