KETriage: Incident Triage with Evolving Hierarchical Symbolic Knowledge
Abstract
Cloud incidents can degrade user experience and cause financial losses, while identifying the responsible service through incident triage remains challenging due to large blast radii and complex fault propagation. Despite extensive efforts, existing approaches still face two fundamental challenges: the absence of integration mechanisms for domain knowledge and brittle adaptation to evolving systems. We present KETriage, a knowledge-enhanced and self-improving agentic framework that identifies the responsible service. To bridge the knowledge gap of LLM agents in incident triage, KETriage constructs a hierarchical symbolic knowledge base (HSKB) offline, a symbolic graph organized into ontology, instance, and case layers. During online inference, KETriage anchors incident terminologies to the HSKB ontology taxonomy and retrieves relevant knowledge segments to identify the responsible service. Given the inherent dynamic nature of cloud systems, KETriage further introduces post-resolution knowledge consolidation, enabling continual experience evolution without prohibitive retraining. We conducted comprehensive experiments using real-world incidents collected from an industrial cloud service system serving millions of users. The experimental results show that KETriage outperforms existing methods, achieving Acc@1 of 79.48% and 93.66% Acc@3 in incident triage.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.