Policy-driven Automated Red-Teaming Framework for Large Language Models
Abstract
Automated red-teaming is essential for LLM safety, yet existing methods cover a subset of the end-to-end steps needed to go from policy definition to actual attack and evaluation, making systematic comparison difficult. We introduce PARF, a structured framework for defining automated red-teaming methods across four stages: attacker’s objective generation, prompt generation, attack construction, and response assessment. PARF provides standardized components that let any method be inserted at its stage and evaluated while the remaining stages are held fixed. We compare single-turn, multi-turn, and graph-based methods across five safety policies and three target models along three complementary dimensions: attack success rate, measuring effectiveness; prompt diversity, measuring coverage of the vulnerability surface; and token cost, measuring practical feasibility, a dimension often omitted in red-teaming studies despite the high expense of these methods at scale. Because PARF abstracts the full red-teaming workflow, it enables us to measure prompt diversity across methods under identical conditions, revealing that methods exploit non-overlapping vulnerability surfaces: many successful attacks are reachable by only one approach, making the combination of methods essential for expanding coverage.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.