acceptodds
Under review as a conference paper at ICLR 2027

RedLoop: Maintaining Reusable Attack Strategies for Multi-Round LLM Safety Training

Abstract

Multi-round safety training must reduce harmful answers without encouraging refusal of benign queries. As training changes the defender, the attack strategies a red-teamer has stored can stop working, and its library can miss strategies that would expose the updated defender; retraining the red-teamer or rechecking every stored strategy is costly. We present RedLoop, a multi-round safety-training framework that keeps the attack-generation model frozen and adapts the red-teamer, RedScout, through a persistent library of reusable attack strategies, natural-language instructions for generating adversarial prompts. To keep the library useful without rechecking every strategy, *fitness propagation* estimates how well previous-round strategies still work from a sample of the current round's attacks and reuses the highest-ranked ones in the next training batch. Alongside strategies distilled from successful attacks, *novelty injection* contributes strategies that are distinct from the library and succeed against the defender, and later attack steps draw on them more heavily. The defender is trained with GRPO under a reward that penalizes refusal of benign queries and of benign queries carrying adversarial framing. On Nemotron-Nano-4B, RedLoop reduces HarmBench attack success from % to and XSTest over-refusal from % to %, while IFEval and MMLU stay within one point of the base model; removing fitness propagation raises HarmBench attack success to % and removing novelty injection raises XSTest over-refusal to %. Safety gains also extend to Llama-3.1-8B and OLMO-3-7B defenders.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.