SuperRed: AI Red-Teaming and Security Benchmarking with a Fine-grained Threat Model
Abstract
Current AI red-teaming evaluations are fragmented across standalone attacks and benchmarks with different threat models, making reported attack success rates (ASRs) difficult to compare and reproduce. For the same attack strategy or benchmark, changing the attacker model and access scope moves ASR from 2.7% to 67.0%. We propose the open-source framework SuperRed for modularized AI red-teaming and security evaluation. SuperRed explicitly specifies and enforces attacker capabilities, including access, feedback, model, and budget, under a common execution model. Attackers, targets, and benchmarks are interchangeable modules, enabling compatible combinations to be evaluated without strategy-specific integration. Using SuperRed, we evaluate compatible combinations of 11 attack strategies, 6 benchmarks, 13 attacker models, and 11 access scopes, using approximately USD 130,000 of inference across chatbot and agent targets. Open-weight models are the strongest tested jailbreak attackers and strong against agent targets (Claude Code). Observing target trajectories provides the largest access advantage; together with judge feedback, it raises AgentDojo ASR from 15.2% to 58.9%. Increasing attacker budget shows sharply diminishing returns after USD 0.19 per jailbreak task and USD 0.25 to USD 0.57 per agent task. These results show that attack success is conditional on the threat model and should be reported and compared accordingly, and that a poorly resourced attacker with an open-weight model already attacks close to the ceiling of current strategies.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.