Semi-Supervised Tabular Anomaly Detection with Contamination-Tolerant In-Context Learning
Abstract
In-context learning (ICL) enables pretrained tabular models to adapt to new tasks without task-specific parameter updates by treating training instances as context. However, directly applying ICL to semi-supervised tabular anomaly detection is challenging because unlabeled contextual data may be contaminated by hidden anomalies, while only a few labeled anomalies are available to provide reliable supervision. Moreover, open-world unlabeled data may contain unknown anomaly types that are not represented by the labeled anomalies. To address these challenges, we propose TabPFN-AD, a framework that iteratively constructs a more accurate in-context set under the semi-supervised tabular anomaly detection setting. TabPFN-AD leverages the inter-instance attention of a pretrained tabular foundation model to evaluate the degree of anomalousness at both the instance and group levels. By mining both known and unknown anomalies from unlabeled data, it maximizes the utility of limited labeled supervision to improve detection performance. Experimental results demonstrate that TabPFN-AD achieves state-of-the-art performance on 17 benchmark datasets and exhibits strong robustness under scenarios where unlabeled data is mixed with a high proportion of potential anomalies and labeled anomalies are extremely scarce.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.