FADE: Post-Hoc Failure Detection and Uncertainty Estimation Across Classical and Vision Foundation Models
Abstract
Vision Foundation Models (VFMs) have set new standards in zero-shot and open-vocabulary segmentation and object detection. Yet reliably assessing their predictions remains a central challenge: model-reported confidence scores are often poorly calibrated. We introduce FADE, a generic post-hoc framework that trains meta-classifiers to discriminate correct from incorrect predictions based on a predefined IoU threshold, using a common representation built from model-reported confidence, geometric features, and semantic information from a frozen vision-language encoder (CLIP). Requiring neither modifications to the underlying model nor access to its internal representations, FADE is fully model-agnostic and applies to classical as well as modern VFMs, spanning both segmentation and object detection. We use FADE to systematically compare prediction reliability between classical and foundation models across multiple datasets and tasks. Our results demonstrate that FADE reliably estimates the probability of success for both model types using this model-agnostic representation. Moreover, in terms of Expected Calibration Error (ECE), estimating this probability of success is harder for foundation models than for classical models. This points to a concrete limitation in the reliability of current VFMs, which a trustworthy predictor should not exhibit. Our code is publicly available at TBA.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.