When Does Diversity Help? Realizable Complementarity in Neural Operator Ensembles
Abstract
Neural operators are commonly developed and evaluated as individual predictors, with improvements largely pursued through architectural advances. Yet our experiments reveal a surprising result: a simple average of independently trained models can deliver substantial accuracy gains. Across six PDE tasks, simple equal-weight averaging reduces error by up to , and the effect persists across different model families and controlled settings. However, the benefit varies sharply across expert pools, revealing that predictive diversity alone does not imply useful aggregation. We show that what matters is whether differences among models produce complementary errors that actually improve the combined prediction. This motivates the notion of useful complementarity. We further uncover a second gap: useful structure may exist within an expert pool without being reliably exploitable from finite development data. Increasing the flexibility of the combination rule does not necessarily translate into larger held-out gains, indicating a distinction between complementarity that is available and complementarity that can actually be realized. We call the latter realizable complementarity. Finally, our results show that information useful for constructing a strong aggregate need not remain equally useful for refining it afterwards. Together, these findings provide a framework for understanding when diversity among neural operators translates into reliable accuracy gains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.