Modality-Agnostic Anomaly Detection with Deterministic, Decomposable Uncertainty Quantification
Abstract
Anomaly detection (AD) increasingly relies on multi-modal data, yet existing methods use fixed-modality architectures, suffer domain interference from shared reconstruction paths, and output only deterministic predictions without reliable uncertainty. We propose a unified framework combining a Multi-Modal Feature Extraction Backbone (MM-FEB) and a Deterministic Uncertainty Quantification Network (DUQNet). MM-FEB unifies arbitrary modality combinations (RGB, depth, infrared) into a shared feature space and applies a hierarchical multi-scale compression structure that suppresses anomaly-related noise while preserving normal patterns, mitigating cross-modal interference. DUQNet extracts region-level features and trains lightweight MLPs to directly regress predictive uncertainty, decomposing it into three interpretable sources: epistemic model uncertainty, aleatoric proposal uncertainty, and task uncertainty, without requiring calibration data or sampling. Experiments on nine public datasets spanning industrial, synthetic, and medical domains show that our method achieves 99.089% on MVTec-3D, 96.913% on Eyecandies, and 97.468% with a state-of-the-art 51.404% on the challenging BraTs benchmark, where the low reflects the small and ill-defined tumor boundaries rather than weak localization. Our method consistently outperforms state-of-the-art detectors, while reducing the pixel-level Expected Calibration Error (ECE) to 0.049 on BraTs and 0.054 on MVTec-3D, corresponding to an average 37.3% relative reduction compared with the best competing method. These results support practical applications such as adaptive model selection, input quality assessment, and task supervision at low computational cost. The code is available at https://anonymous.4open.science/status/iclr2027-B601.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.