Saturated Evaluation Looping Attacks in Multi-Agent Systems
Abstract
Evaluation-looping schemes are now widely used in multi-agent systems to enhance performance across various tasks. Yet, while designed to improve task effectiveness, we highlight that these schemes also introduce a direct attack surface. Specifically, they can give rise to a previously underexplored attack scenario: the saturated evaluation looping attack, wherein an adversary seeks to deliberately drive a multi-agent system into persistently saturated evaluation loops to substantially prolong system runtime. To systematically uncover the adversarial risks in this scenario, we develop a novel attack method tailored to it. Extensive experiments show that our method effectively exposes the risks in this scenario, offering new insights into the robustness of multi-agent systems. We will release the code.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.