The Target Agent Is Its Own Best Critic—but a Poor Proposer: Role Separation for Automatic Red Teaming
Abstract
Automatic red teaming increasingly relies on large language models (LLMs) to iteratively criticize failed attacks and propose improved adversarial prompts or prompt injections. Although existing methods strengthen this loop through prompt engineering and search-harness design, the mechanisms underlying its two core roles, criticism generation and attack proposal, remain poorly understood. We systematically disentangle two design choices: whether criticism should be generated by the target session itself (self-criticism) or by a separately instantiated context (external criticism), and whether that criticism should be translated into the next attack by the target session (self-proposal) or by a separate attacker (external proposal); we further investigate why these choices lead to different outcomes. Our analysis reveals that self-criticism is more effective than external criticism because it is grounded in the target’s live execution context, including its observations, constraints, and available reasoning; removing reasoning reduces downstream attack success by 6–17 percentage points. In contrast, self-proposal is substantially less effective than external proposal because awareness of the adversarial objective induces both explicit refusal and sandbagging. Based on these findings, we introduce a role-separated red-teaming method that uses the target session to diagnose attack failures and a separately instantiated attacker to generate the next proposal. We evaluate this design on DTAP-BENCH across Telecom, Workflow, CRM, and Code domains, covering direct misuse and indirect prompt injection against agents implemented with multiple frameworks. Our method achieves an overall attack success rate of 63.2%, achieving the state-of-the-art result. These findings show that effective automatic red teaming depends not only on model capability or search strategy, but also on how execution context and adversarial responsibilities are distributed across model instances and context in the red-teaming harness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.