acceptodds
Under review as a conference paper at ICLR 2027

Toward Valid Measures of Efficacy and Success for AI Red-teaming

Abstract

Red-teaming has become a cornerstone of GenAI safety evaluation, but the field lacks principled criteria for measuring its efficacy and success. We identify several core challenges that undermine the validity of existing metrics, such as Attack Success Rate (ASR). To understand and address them, we propose a probabilistic framework that grounds red-teaming success in a single well-defined quantity: the safety risk of the model (or expected severity of its outputs) under a target usage distribution p* capturing the threat model of interest. We show that the gap between this quantity and any empirically computed success metric is controlled by the distance between the red-teaming distribution q and the target distribution p*, the smoothness of the underlying severity function outside a small set of irregular prompts, the validity of severity assessments in practice, and a sampling error that shrinks with corpus size. Our formalization offers three core insights: the validity of red-teaming success metrics depends on 1) the alignment of the attack distribution under red-teaming and under the threat model of interest, 2) the empirical assessment of severity matching the real severity function, and 3) the smoothness of model safety behavior. In particular, because more prompts reduce only the sampling error, validity does not necessarily improve with evaluation corpus size or superficial attack diversity. We apply the framework to two corpora. Holding the target model and safety judge fixed and varying only the evaluation distribution, measured ASR ranges from 0.118 to 0.905, and distance from the reference distribution orders the resulting error exactly. On a corpus of 38,659 human red-teaming conversations we then estimate the smoothness pair directly, reporting what are to our knowledge the first empirical estimates of severity smoothness in this literature and showing that preference-trained models exhibit markedly smoother safety behavior than base models. We discuss the implications of our formalization for evaluation practice and outline better approaches for measuring success going forward.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.