DART: Defense-Aware Red Teaming with a Critical Cue Bank for Open-World Adversarial Evaluation
Abstract
Autonomous red-teaming agents can inspect source code, invoke attack libraries, retrieve external knowledge, and use execution feedback, yet these capabilities do not ensure that the agent can identify relevant knowledge and operationalize it as an executable attack tailored to a particular defense. We introduce DART, a defense-aware open-world red-teaming framework centered on a Critical Cue Bank (CCB). The CCB structures profiling evidence into actionable cues that guide toolbox attack adaptation, external retrieval, and research-to-executable attack construction. When no directly applicable attack is available, DART retrieves strategies developed for related defense mechanisms and adapts their principles to the target's interface and threat-model constraints. Across 51 published defense implementations from AutoAdvExBench, DART increases the overall attack success rate from 5.9% to 43.1% with Qwen and from 31.4% to 52.9% with Claude Opus, while producing valid evaluations for all 51 targets. It also reduces LLM token usage by 84-95% relative to the corresponding baselines. An open-world baseline without cue-conditioned selection and adaptation yields little or no improvement in attack success rate, suggesting that the main bottleneck is not access to attack knowledge but its target-conditioned operationalization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.