A Reassessment of Global Weight-Space Sampling via Deterministic Flat Minima
Abstract
Deep Ensembles remain the empirical upper bound for Uncertainty Quantification, successfully capturing global epistemic diversity across disjoint loss valleys. To circumvent their prohibitive computational overhead, modern frameworks frequently rely on local weight-space sampling (e.g., SWAG, Laplace Approximations) to approximate an ensemble within a single basin. In this paper, we analyze the scaling issues that arise when global weight-space sampling is applied to overparameterized architectures, formalizing this as the weight-space variance dilemma. We highlight a geometric tension: strictly bounding stochastic weight perturbations within a certified optimal-transport safety radius leads to dimensionality dilution, where noise merges with the high-dimensional null space, leaving predictions functionally identical to the base checkpoint. Consequently, practical single-basin samplers are forced to inject unconstrained variance to achieve ensemble diversity. Such macroscopic noise can disrupt internal feature manifolds, necessitating computationally expensive deployment heuristics such as data-dependent batch normalization recalibration. Based on recent theoretical insights that flat minima can implicitly regularize predictive entropy, extensive evaluations on ResNet-50, ViT-B/16, and YOLOv8x demonstrate that much of the out-of-distribution logit smoothing attributed to global sampling can be natively embedded in the deterministic base optimizer. While unable to model true multi-modal posteriors, a single deterministic Sharpness-Aware Minimization (SAM) checkpoint serves as a highly efficient, zero-overhead baseline that frequently matches the local calibration of global ensembles. Furthermore, we show that while uniform Euclidean bounds stabilize homogeneous classifiers, complex dense prediction tasks require scale-aware Mahalanobis-Wasserstein spaces (ASAM) to prevent structural regression collapse.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.