acceptodds
Under review as a conference paper at ICLR 2027

Portfolio of Experts

Abstract

Sparse mixture-of-experts (MoE) models increase model capacity without proportionally increasing per-token computation by activating only a small subset of experts. Route quality therefore depends on both expert utility and complementarity within the selected subset. However, most routers rank experts independently, potentially selecting experts with redundant task contributions. Effective routing must instead identify a high-utility, complementary expert subset under a fixed budget. We introduce Portfolio of Experts (PoE), which formulates sparse routing as conditional portfolio optimisation under outcome uncertainty. PoE defines counterfactual gradient returns (CGRs), which quantify the local effect of each expert's allocation on downstream loss. A lagged critic predicts each candidate expert's expected CGR and the pairwise covariance of these returns. A constrained solver then jointly selects and weights a small set of experts, favouring high predicted utility while avoiding redundant combinations under a fixed routing budget. Across seven zero-shot downstream tasks spanning science question answering, commonsense reasoning, reading comprehension, and contextual prediction, PoE attains the highest mean accuracy and lowest held-out NLL at every evaluated expert-pool size. Frozen-route studies show that modelling pairwise dependence improves local routing utility, while routing diagnostics show that PoE lowers joint failure without activating more experts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.