On the Framing Effect of AI Red-Teaming Instructions: Evidence from Prior Exercises and Controlled Simulations
Abstract
AI red-teaming, adversarially probing generative AI systems to surface harms, has become a central evaluation method. However, prior work has highlighted that red-teaming is a fairly open-ended task, laden with myriad design choices that can shape its findings. In this work, we focus on the instructions used to conduct red-teaming exercises and hypothesize that they can significantly affect outcomes in a manner analogous to the framing effect in psychology, where the format in which information is presented greatly impacts how people respond. To understand whether and why instructions impact red-teaming outcomes, first, we analyze the range of ways in which instructions have historically been formulated, reviewing 228 instances of prior work and extracting thirteen dimensions along which instructions vary in practice. Second, based on takeaways from this review, we run controlled simulations that hold parameters such as the risk domain, the attacker, and the model under investigation fixed while varying only the instructions: from five real instruction templates used by industry and academic organizations, we derive 45 instruction frames across three risk domains and three levels of risk granularity, generating 100 variants per frame, and deploy each against five open-weights target models, for 22,500 attacks. Using a generalized linear mixed model, we find that certain instruction properties significantly predict attack success rates: the lowest and highest levels of risk granularity both outperform intermediate specificity, as do short and long instructions. The presence of examples is associated with higher attack success rates, while supplying explicit harm definitions yields lower ones. We replicate these patterns with a frontier model as attacker and target. As alignment evaluations become increasingly automated, our results suggest that humans continue to shape red-teaming outcomes substantially through the instructions they provide to scaffold the activity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.