CRETAP: From Reconstruction Errors to Reliable Anomaly Probabilities on Imbalanced Data
Abstract
Reconstruction-based anomaly detectors, such as autoencoders, rank samples by an error whose scale depends on the dataset, preprocessing, and aggregation, so thresholds do not transfer and scores are not probabilities. Fixing this with post-hoc calibration withholds scarce labeled anomalies from training, and switching to supervised classifiers discards the element-wise residuals that explain each decision. We introduce CRETAP, which fits Platt scaling of the log-error—a two-parameter Hill function—jointly with the reconstruction network under binary cross-entropy (BCE) instead of afterwards, so the anomalies that train the detector also calibrate it. Because the probability depends on the input only through the error, we prove that the excess BCE upper-bounds twice the squared calibration error and that stationary points of training are mean-calibrated with unit calibration slope in-sample. The forward pass also yields feature-level attributions—each element's error contribution divided by the learned decision boundary—that sum to more than one exactly when the probability exceeds one half. On the 29 ADBench datasets with at most 5% anomalies and without any post-hoc calibration, CRETAP is the best-calibrated of eight natively probabilistic detectors on 9 datasets, more than any other, and second in mean expected calibration error (0.0073 vs. 0.0067 for XGBOD), while its mean area under the precision–recall curve, 71.19%, trails CatBoost's by 2.0 points. In this regime, a reconstruction detector can thus serve as its own calibrator; above 5% contamination, gradient-boosted trees are preferable.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.