RedOpus: Multi-turn Red-teaming LLM Agents with Reflective Memories
Abstract
The development of Large Language Model (LLM) Agents that can invoke tools and act upon external data and entities is built upon LLMs' instruction-following and reasoning capabilities. The insufficient robustness in these capabilities, however, raises security risks and facilitates red teaming research aiming to uncover agent vulnerabilities under complex, noisy and sometimes adversarial conditions. While existing works focus on one-shot attacks and defenses, real application scenarios usually involve multi-round interaction where LLM agents face continuous, strategic threats. To bridge this gap, we propose RedOpus, a multi-turn red-teaming framework that leverages reflective memory to guide iterative attack refinement. We also propose RedOpusArena, a multi-turn interactive sandbox featuring diverse simulated scenarios and automatic generation of massive tasks. We test our red-teaming framework on agents with multiple LLM backbones and reveal a general vulnerability of existing LLMs to automatic iterative attacks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.