acceptodds
Under review as a conference paper at ICLR 2027

Do Deep Ensembles Actually Capture Uncertainty in Graph Neural Networks?

Abstract

While deep ensembles are widely considered to be the default method for uncertainty quantification in deep learning, their effectiveness for graph-structured data is often simply assumed based on successes in domains like computer vision. We investigate standard deep ensembles specifically for message-passing graph neural networks. Benchmarking across eight datasets representing varied tasks and complexities, we reveal that ensembles provide only marginal improvement over a single model. These marginal gains stem primarily from averaging out optimization noise in point predictions rather than from meaningfully better uncertainty estimates. Through an aleatoric-epistemic decomposition, we identify epistemic collapse: independently trained networks consistently converge to overly similar predictions. Because disagreement is the fundamental mechanism through which ensembles capture epistemic uncertainty, this lack of diversity neutralizes their key advantage. Towards explaining this phenomenon, we derive a bound on epistemic variance in terms of the graph topology and Dirichlet energy budget, linking epistemic collapse to the smoothing bias of message passing. Our results suggest that deep ensemble success does not seamlessly transfer to graph machine learning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.