Cut My Pool Into Pieces: Rethinking Model Pool Construction for LLM Routing
Abstract
The enormous popularity of large language models (LLMs) has led to a vast and highly diverse landscape of models, varying in their strengths, weaknesses, and inference costs. As different models are suitable for different queries, LLM routing is increasingly applied to select an appropriate model for each query while balancing response quality and cost. Nonetheless, a router can only select among the models available in its pool, so the composition of that pool directly affects the routing system's quality-cost trade-off. Despite this, prior work has primarily focused on the router itself, paying considerably less attention to how the model pool should be constructed. In this work, we study model pool composition through pruning, systematically evaluating 10 pruning strategies across 27 benchmark tasks, 39 distinct models, and six router families. We find that well-pruned pools can preserve nearly all of the full pool's routing performance, while pruning strategies vary widely in their ability to serve low-cost budgets. Pool composition is particularly important for small pools, where its impact on routing performance exceeds that of router choice itself. Finally, pruning transfers well to unseen tasks, with performance close to that of in-task pruning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.