An Empirical Evaluation Framework for Auditing Explainable AI in Network Intrusion Detection
Abstract
Deep learning Network Intrusion Detection Systems (NIDS) suffer from a transparency gap that hinders operational trust. While Explainable AI (XAI) methods aim to bridge this, they often exhibit domain blindness, generating semantically invalid network states, such as fractional TCP flags that distort evaluation metrics. This paper proposes a modular evaluation framework centered on a novel Protocol-Aware XAI (PA-XAI) engine. By enforcing protocol invariants and hierarchical constraints, PA-XAI ensures that explanations are grounded in physically realizable telemetry. Our framework employs four auditing modules: functional faithfulness, adversarial robustness, algorithmic consensus, and dimensionality reduction utility. Empirical results across four benchmark datasets reveal that standard XAI metrics often mask significant instability; for example, LIME’s apparent stability is frequently an artifact of invalid neighborhood construction. Conversely, our PA-variants, specifically PA-DeepLIFT and PA-IG, demonstrate superior robustness against adversarial evasion. Furthermore, PA-XAI guides dimensionality reduction to achieve up to 90.5% feature compression with less than a 1% drop in detection fidelity. These findings demonstrate that domain-aware auditing is essential for transitioning XAI from a theoretical enhancement to a reliable component of operational network defense.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.