Exploring a Principled Framework for Multi-label Out-of-distribution Detection with Pre-trained Vision-Language Models
Abstract
Out-of-distribution (OOD) detection enhances the reliability of machine learning models by identifying unexpected inputs from unknown classes. Recent progresses in large-scale vision–language pre-training have enabled zero-shot OOD detection using only in-distribution (ID) class names. Nevertheless, previous studies predominantly focus on detecting OOD inputs in the multi-class classification task, leaving the more realistic multi-label setting greatly under-explored. To mitigate this research gap, this paper investigates zero-shot multi-label OOD detection with pre-trained vision-language models. We first theoretically demonstrate that existing methods are fundamentally mismatched with the multi-label setting, as their scoring mechanisms fail to effectively aggregate evidence from multiple co-occurring ID labels. Motivated by this, we propose a probabilistic framework that formulates multi-label OOD detection through pairwise image-label association modeling, directly yielding a principled OOD scoring function that estimates the probability of an input being associated with at least one ID label in a training-free manner. Comprehensive evaluations on standard multi-label benchmarks, including MS-COCO and PASCAL VOC, empirically manifest that our method achieves state-of-the-art performance across a variety of OOD detection setups.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.