Information-Theoretic Query Complexity for LLM Ensemble Selection
Abstract
Selecting an LLM ensemble is fundamentally about complementarity: a model is valuable when it succeeds where others fail. We study how many evaluations are needed to identify the best fixed-size ensemble and show that comparison difficulty is determined entirely by disagreement. Tasks that both ensembles solve or both miss are uninformative; only outcomes on which exactly one succeeds matter. For binary feedback indicating whether each model succeeds, this yields the exact information rate governing reliable comparison and, when all models are evaluated on the same sampled tasks, the exact asymptotic number of sampled tasks required to identify the best ensemble. The same principle extends to model-ranking feedback. Building on this, we develop a greedy procedure that is asymptotically optimal in the number of accepted tasks for identifying the model that best solves tasks missed by the current ensemble. Yet, this greedy decision can require arbitrarily more sampled tasks under full-vector feedback than identifying the best ensemble directly, even when greedy ultimately selects it. Direct identification faces a separate computational limitation: even when the required number of sampled tasks is polynomially bounded, computing the exact best ensemble in time polynomial in that bound and the number of models on every instance would imply . Experiments on published LLM evaluations and a fresh multilingual benchmark are consistent with these predictions. Even with the same overall success gap, ensembles can require substantially different numbers of sampled tasks to distinguish depending on how their successes overlap. Exploiting this complementarity can outperform selection by individual accuracy. Compared with the closest existing method, ours requires less evidence to confirm its choices while achieving similar average ensemble quality. Finally, realistic benchmarks can enable improved ensemble selection while containing too few tasks to confirm the choices with high confidence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.