acceptodds
Under review as a conference paper at ICLR 2027

ARMOR: Agentic Red-Teaming with Memory-Optimized Retrieval for Autonomous Jailbreaking of Frontier Language Models

Abstract

Automated red-teaming provides a systematic approach for identifying safety vulnerabilities in large language models (LLMs) prior to deployment. Existing black-box methods, such as PAIR and TAP, rely on stateless iterative prompt refinement, limiting their ability to retain and reuse successful attack strategies across interactions. We introduce ARMOR, a memory-augmented multi-agent red-teaming framework that addresses this limitation through persistent strategy retrieval. ARMOR coordinates a Strategist agent, an Attacker agent, and a three-judge Jury panel with an Arbiter for tie-breaking, orchestrated through a LangGraph state machine. Successful attack strategies are stored in a vector database and retrieved to initialize subsequent attacks, enabling the system to accumulate and reuse adversarial knowledge across interactions. We evaluate ARMOR across five model tiers spanning 8B to 550B parameters and two frontier model families, LLaMA and NVIDIA Nemotron. To reduce evaluator circularity, successful attacks are independently cross-validated using Meta's Llama Guard 4 (12B) moderation classifier. We further introduce Expected Target Queries (ETQ), a metric designed to account for survivorship bias in baseline efficiency estimates. ARMOR requires 3.17 ETQ on LLaMA-3.3-70B, compared with 15.69 for PAIR and 11.19 for TAP, while achieving a 96% raw attack success rate (ASR). On the heavily aligned Nemotron-120B, ARMOR achieves 86% raw ASR, compared with 48% for PAIR and 17% for TAP. Independent validation further shows that only 38.37% of the raw successes produced by PAIR and TAP are classified as genuinely unsafe by Llama Guard, compared with a 52.99% overall agreement rate for ARMOR. Ablation studies show that persistent memory retrieval improves ASR by 43 percentage points on Nemotron-120B relative to the memory-ablated variant. Finally, ARMOR achieves 20% ASR against a 550B-parameter model in a pure black-box setting, compared with 5% for stateless baselines, demonstrating the utility of persistent adversarial memory for automated safety auditing of frontier LLMs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.