TRIAD: Calibrating Contextual and Normal-Support Evidence for Unsupervised Multi-Class Anomaly Detection
Abstract
Unified multi-class anomaly detection asks a single model, trained only on normal images, to detect and localize defects across many product categories. Scoring pipelines that inspect the target they are scoring suffer from two structural failures. First, attention lets target content inform its own score, so a plausible defect can be explained as normal. Second, raw error maps lack a common probabilistic scale across positions, appearance modes, and categories. We present TRIAD (Target-blind energy, Reference witness, and Integrated Anomaly Detection), a framework built entirely on a frozen DINOv2 backbone that computes three calibrated evidences for every test image. A target-blind contextual energy fills in masked patches from context alone and grades the observation by its likelihood under a von Mises–Fisher mixture over a library of normal patch atoms. A witness branch measures the distance to a coverage-maximizing library of normal images, aggregates the distance with a CVaR tail readout, and converts the aggregate into route-conditional empirical -values with finite-sample false-alarm control. A calibrated fusion places both fields on one commensurable scale and combines them with equal weights, a statistic that is likelihood-ratio optimal for dense shifts under a Gaussian working model, while image-level decisions remain identical to the witness detector. On MVTec-AD and VisA under the unified protocol, TRIAD advances the state of the art on all seven reported metrics, reaching / pixel AUROC, / pixel AP, and / AUPRO.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.