BoundaryPlay: Improving Online Safety Self-Play via Counterfactual Request Construction
Abstract
In a weak-safety setting, online safety self-play adapts training requests to a changing defender, but how harmful and benign requests are constructed remains underexplored. Independent generation leaves their relationship unconstrained. We introduce BoundaryPlay, a counterfactual request-construction protocol: the attacker uses a seed goal and shared framing, aiming to vary safety intent between harmful and legitimate benign requests while preserving unrelated content. The pair is scored jointly for inducing harmful compliance and benign refusal, and the attacker receives feedback on observed defender weaknesses; the per-sample defender objective is unchanged. Against an independent-generation control matched in generated fraction and defender candidate exposure, BoundaryPlay improves the JailbreakBench joint safety–compliance score on Llama-3.1-8B-Instruct-abliterated by 6.58 points (4.45 over a Self-RedTeam RL-only reproduction), with gains positive in all six seeds and significant in a post-hoc Holm sensitivity analysis, and reduces HarmBench harmful responses by 5.69 points against the same control, at a 3.41-point cost in FalseReject compliance. On initially aligned Llama, we identify the defender's plus-side refusal-correctness reward as a modifiable contributor to instruction-following loss: removing this component while retaining counterfactual construction recovers 11.83 points of Instruction-Following Evaluation (IFEval) strict prompt accuracy and improves all four benign compliance panels, with few additional harmful responses on the evaluated harmful-request panels. BoundaryPlay thus contributes both a request-construction protocol and a targeted reward-design intervention that mitigates instruction-following loss in the tested single-seed setting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.