Inference-Time Model Selection via Capability Synthesis
Abstract
We consider the problem of inference-time model selection in language models, where a selector has access to a finite pool of models and aims to select the model that performs best on a given task. In this setup, querying every model on each task is computationally expensive; therefore, the goal is to design contextual model selection strategies that can predict the best-performing model for each context. We propose a simple framework named *Capability Synthesis* that discovers directions along which the models' capabilities are most distinguishable. Capability Synthesis learns to jointly represent models and tasks in a shared capability space, and uses these representations to predict the performance of each model on any given task at inference time. This framework is particularly suited for model selection over large pools of models, as larger pools provide more observations for learning the capability structure, while the cost of the selection strategy remains negligible at inference time. Finally, our experiments on multiple routing benchmarks show that this approach, which admits closed-form solutions, competes with more complex trainable routers and outperforms single-model deployments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.