Beyond Predictive Criteria: Optimization and Downstream Uncertainty
Abstract
Optimization choices are typically guided by predictive criteria such as accuracy, generalization, and calibration, yet many downstream uncertainty estimators operate on learned representations. We ask how uncertainty readouts respond to representations produced by predictively favorable optimization changes. Using the perturbation radius of Sharpness-Aware Minimization (SAM) as a controlled training path, we find a selective divergence: accuracy and calibration improve, output-based OOD scores remain stable or improve, and kNN improves, whereas covariance-sensitive Mahalanobis scoring and Deep Deterministic Uncertainty (DDU) can deteriorate at large SAM radii. We investigate this degradation using Mahalanobis scoring as a tractable case. Within each checkpoint, we vary the strength of covariance weighting while holding the learned representation, class centroids, and covariance eigensystem fixed. Under large-radius SAM, stronger inverse-covariance weighting can become substantially detrimental to OOD detection. Directional analysis links this behavior to a covariance–contrast mismatch in which later low-variance eigendirections receive increasing emphasis despite providing weaker ID–OOD contrast. Crucially, an unweighted readout over the same representation retains strong OOD discrimination as these directions are added, showing that the degradation does not require a uniform loss of OOD-discriminative structure. This covariance-sensitive degradation under large-radius SAM is not universal across sharpness-aware optimizers or ID training distributions, while the broader separation between predictive improvement and representation-dependent uncertainty also appears under SWA and in a ViT/ImageNet-200 setting. Together, these results show that uncertainty readouts can respond differently across representations produced under different optimization conditions and, in the Mahalanobis case, that detector degradation can arise from detrimental inverse-covariance weighting even when useful OOD-discriminative structure remains accessible.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.