acceptodds
Under review as a conference paper at ICLR 2027

LLM Safety: What Are Red Teams and Blue Teams Really Competing For?

Abstract

Large language models (LLMs) are trained to refuse harmful requests, yet carefully crafted jailbreak prompts continue to bypass these safeguards. The standard metric for evaluating attack effectiveness is the Attack Success Rate (ASR). However, we demonstrate that ASR is systematically misleading: it primarily reflects the agreement between the supervisor used to guide prompt generation and the evaluator used to judge success, rather than the intrinsic harmfulness of the model's output. This supervisor-evaluator gap inflates reported success rates, as many attacks that score positively simply exploit shared biases between the two classifiers without eliciting genuinely harmful content. Compounding this issue, the common practice of optimizing for refusal avoidance induces semantic drift, where a malicious query such as “how to make a bomb” is mutated into a benign variant like “how to make a bomb chicken” that evades refusal but produces harmless output. We argue that jailbreak research should shift its objective from making the model say yes to making the model produce actually harmful content. This reframes the problem as an arms race: the red team must learn the blue team's safety boundary to identify blind spots, while the blue team must expand its coverage to close them. We formalize this dynamic through the lens of harm discriminators, showing that attack success depends on the red team's ability to construct a discriminator with high recall relative to false positive rate, and to approximate the blue team's discriminator in black-box settings. Our experiments demonstrate that when the red team directly leverages the blue team's own evaluator as its supervision signal, it discovers truly effective jailbreaks substantially faster. This work offers a fresh perspective on LLM safety testing: the decisive factor is not the sophistication of attack algorithms, but the ability to learn the opponent's hidden decision boundary first.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.