acceptodds
Under review as a conference paper at ICLR 2027

Unified Hierarchical Red-Teaming: Strategy Search and Evidence-Guided Policy Evolution

Abstract

Large language models (LLMs) remain vulnerable to adversarial attacks after safety alignment, motivating automated red-teaming methods that systematically search for model failures. Existing strategy-based methods are mostly designed for a specific attack task and often use final attack outcomes to improve different attack decisions without distinguishing where the failure occurs. We propose a unified red-teaming framework consisting of a Unified Hierarchical Red-Teaming Policy and Hierarchical Evidence-Guided Policy Evolution (HEPE). The hierarchical policy separates task-specific attack knowledge from a shared decision process: each task maintains its own strategy library, interaction protocol, and evaluator, while the shared policy performs strategy selection, structured program construction, and response-conditioned action generation. A two-level tree search explores multiple strategy branches and executes each strategy through a structured attack program and sibling candidate actions under a shared target-query budget. HEPE improves the shared policy using decision-specific evidence collected during search. Same-instance strategy comparisons revise strategy selection, program and execution records revise program construction, and same-state sibling comparisons revise action generation. Compatible evidence for these shared components is aggregated across tasks, while each task-specific strategy library is revised only from cross-instance evidence collected for that task through structured ADD, REFINE, and MERGE operations. Candidate updates are evaluated against the current version on held-out instances using paired validation before acceptance. We evaluate the framework on jailbreaking (JB), system-prompt extraction (SPE), and direct prompt injection (DPI) across six target models. The results show consistent improvements over recent automated red-teaming baselines across all three tasks under limited target-query budgets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.