Which Models to Mix Depends on How You Mix Them: Composition-Aware Portfolio Selection for Mixture-of-Models
Abstract
Given a large pool of candidate language models, which small subset should be composed into a Mixture-of-Models? Existing rules score candidates by individual accuracy, error diversity, or complementary coverage, and implicitly value a pool by the probability that at least one member is correct. This assumption fails for an entire class of deployed systems. We distinguish *selection*-type composers, which output a member's answer, from *assembly*-type composers, which synthesize an answer from member evidence, and show that the two induce structurally different portfolio objectives. Under a subgoal decomposition of difficulty, the gap between their ceilings equals the probability that the members solve a query collectively while none solves it individually. This yields a dichotomy: the ceilings coincide exactly when this event is impossible; otherwise, the optimal portfolios can be disjoint, and using either in place of the other incurs a constant loss. Tractability is similarly asymmetric: the selection objective is monotone submodular at every information level, whereas the assembly objective exhibits increasing returns and admits no polynomial-time constant-factor approximation unless P = NP. Therefore, we recover a guarantee through a submodular logarithmic surrogate. Graders observe competence only after a member's upstream errors have propagated, so we build the surrogate on intrinsic competence, recovered in closed form by inversion over the prerequisite order. Without it, a selector can discard specialists that make assembly worthwhile. Our method performs best on six benchmarks with a 42-model pool, and its diagnostic indicates whether a homogeneous or heterogeneous portfolio performs better before composition runs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.