Deployment-Time Mixture Optimization for Large Language Models
Abstract
Domain-specific fine-tuning produces language models with complementary strengths, yet an unseen deployment distribution may not be well matched by any single model. We introduce Hierarchical Routing, a mixture optimization method that uses an unlabeled batch to identify the Top-\(H\) most relevant sources and optimizes a weight-space combination of their selected models using autoregressive loss. The hierarchy reduces optimization from the full model bank to a low-dimensional combination of relevant candidates, while preserving complementary specializations. We instantiate the bank with models spanning source domains and robustness levels and evaluate on nine Pile domains and deployment streams whose distributions change across batches. The experiments show that routing is beneficial under heterogeneous and changing conditions, despite higher deployment cost, and that applying test-time adaptation after routing further improves performance. The results show that adaptation to unseen distributions can be solved as navigation and composition in a precomputed model space.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.