acceptodds
Under review as a conference paper at ICLR 2027

Population-Based Self-Play for LLM Safety Training

Abstract

While self-play promises unbounded synthetic data for LLM post-training, standard single-policy methods quickly plateau. Training against only the latest opponent causes policies to cycle between behaviors and collapse onto narrow strategies. We address this by applying population-based self-play via Policy-Space Response Oracles to attacker–defender safety training. In each generation, a new policy is initialized from the base model and trained against the Nash equilibrium of the entire opposing population. To make population payoff matrices scalable at 8B parameters, we introduce a racing evaluation budget and a rollout cache. On standard safety benchmarks, our defender achieves 96.2% average safety, outperforming baseline self-play methods (at most 86.4%) at helpfulness comparable to the other self-play methods and with little loss of general capability. Ablations confirm that each element is necessary: training against single opponents induces cycling, uniform weighting degrades safety, and warm-starting triggers policy collapse. Finally, we show that strong populations expose reward specification flaws, enabling targeted fixes to the reward function.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.