acceptodds
Under review as a conference paper at ICLR 2027

Automated Discovery of Failure Modes in Agentic AI Systems

Abstract

Multi-agent systems composed of large language model-driven agents are increasingly used in real-world applications, yet evaluating their reliability remains challenging. Compared with single-model systems, their decentralized, multi-step, and stochastic behavior complicates failure analysis. We present a framework for automated identification of failure modes in agentic AI systems with limited or no historical interaction data. Our framework generates domain-specific evaluation scenarios under structured categorization of statistical tests to uncover system weaknesses. Our framework is evaluated against a real-world multi-agent customer-service system, enabling realistic failure-mode detection. The results have shown that our framework identifies failure modes missed by state-of-the-art evaluation approaches.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.