acceptodds
Under review as a conference paper at ICLR 2027

Scaling Laws of Ensemble Uncertainty in Language Models

Abstract

Many use cases for large language models (LLMs), from AI safety to exploration and human interaction, require estimates of epistemic uncertainty. While ensembles are a practical option, multi-seed experiments show that their disagreement is small compared to individual token entropy and significantly underestimates the model's true epistemic uncertainty. We try to shed light on these observations and broader aspects of ensembling for uncertainty estimation in LLMs. The upshot of our findings is that ensembling becomes more profitable both for uncertainty estimation and as an alternative to scaling model size at larger scales. We relate epistemic uncertainty to reducible loss in calibrated models, separating the portion visible through ensemble disagreement from an invisible portion that includes model misspecification. We show that the visible portion decreases as a mild power law with model size and changes little with additional training data or context beyond an initial stabilization range. The invisible portion decreases faster, so ensembles reveal a growing share of average epistemic uncertainty as models scale. These scaling trends also predict that ensembling becomes more competitive in terms of loss as the size of a single model increases, with matched total parameters and fixed training data. We then identify a second power law for ensembles: relative to independently trained ensembles, that relates the ensemble diversity to the shared portion of the training across ensemble members. This implies that late branching is heavily penalized in diversity, and the issue is even more severe when ensembling with low-rank adapters.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.