acceptodds
Under review as a conference paper at ICLR 2027

Evidence-Aware Robust Ensembling for Calibrated Uncertainty in Fine-Tuned Foundation Models

Abstract

As foundation models are increasingly deployed in high-stakes domains, their tendency to generate incorrect or low-quality content has become a major reliability concern. This risk is often amplified after fine-tuning for downstream tasks, where domain shift and limited task-specific data can degrade calibration and output quality. A common approach to uncertainty estimation is to sample multiple candidate responses through probabilistic decoding and use their disagreement as an uncertainty signal. From a Bayesian perspective, aggregating multiple generations can be interpreted as approximating marginalization over a posterior distribution over model parameters, whereas selecting a single maximum-probability output corresponds to a point estimate. However, fine-tuning on small or heterogeneous datasets can induce complex, flat, and potentially multimodal posteriors over model parameters, for which decoding-based sampling provides only limited and locally biased coverage, leading to miscalibrated uncertainty estimates. To address this limitation, we propose a diversity-inducing ensemble framework guided by Distributionally Robust Optimization (DRO) and evidential theory. It promotes complementary hypotheses that better cover disparate modes of the fine-tuned posterior, while inducing structured perturbations in the loss landscape to approximate Bayesian marginalization. Meanwhile, evidential modeling enables a principled decomposition of predictive uncertainty. We establish theoretical guarantees linking coverage of complementary posterior modes to improved uncertainty calibration. Experiments on benchmark datasets and large vision-language models show that our method consistently improves calibration and more accurately estimates content quality (correctness of generated outputs through confidence-based filtering), thereby enhancing the reliability of fine-tuned foundation models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.