Automated Discovery of Failure Modes in Agentic AI Systems
Abstract
Multi-agent systems composed of large language model-driven agents are increasingly used in real-world applications, yet evaluating their reliability remains challenging. Compared with single-model systems, their decentralized, multi-step, and stochastic behavior complicates failure analysis. We present a framework for automated identification of failure modes in agentic AI systems with limited or no historical interaction data. Our framework generates domain-specific evaluation scenarios under structured categorization of statistical tests to uncover system weaknesses. Our framework is evaluated against a real-world multi-agent customer-service system, enabling realistic failure-mode detection. The results have shown that our framework identifies failure modes missed by state-of-the-art evaluation approaches.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.