When Does Reference Purification Help? Contamination-Aware Representation Learning for Unsupervised Anomaly Detection
Abstract
Unsupervised anomaly detectors often learn representations and reference distributions from unlabeled data assumed to be predominantly normal. Unknown contamination can distort both the learned geometry and the reference bank used for scoring, but removing suspected anomalies can itself discard useful normal structure. We study when reference purification helps and when it does not. We introduce CI-DCUAD, which combines representation learning with contamination-invariant reference selection (CI-RS). Preliminary unsupervised scores select observations for representation warmup and final reference construction, while main-stage training returns to the complete unlabeled dataset. We characterize effective contamination and show that selection reduces anomalous reference mass exactly when normal observations are retained more frequently than anomalies, . Purification alone, however, does not guarantee better detection, which also requires preserving normal support and informative reference geometry. A controlled separability–contamination study comprising 840 runs identifies distinct failure, success, and abstention regimes. Purification is not practically reliable at zero separation, is reliable in every valid nonempty positive-separation condition, and yields empty selections in two extreme conditions. Across five tabular benchmarks, CI-selected reference construction accounts for most of the favorable F1 change on CreditCard2013 and NSL-KDD, but produces unfavorable effects on the remaining datasets. Neither reference-only nor CI-Full significantly exceeds the matched baseline over the complete structural-ablation grid after Holm correction, although CI-Full significantly exceeds both incomplete CI variants. These results show that reference purification is a conditional mechanism, not a universal robustness strategy, and characterize the conditions under which it succeeds, fails, or should abstain.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.