acceptodds
Under review as a conference paper at ICLR 2027

Deployment-Time Mixture Optimization for Large Language Models

Abstract

Domain-specific fine-tuning produces language models with complementary strengths, yet an unseen deployment distribution may not be well matched by any single model. We introduce Hierarchical Routing, a mixture optimization method that uses an unlabeled batch to identify the Top-\(H\) most relevant sources and optimizes a weight-space combination of their selected models using autoregressive loss. The hierarchy reduces optimization from the full model bank to a low-dimensional combination of relevant candidates, while preserving complementary specializations. We instantiate the bank with models spanning source domains and robustness levels and evaluate on nine Pile domains and deployment streams whose distributions change across batches. The experiments show that routing is beneficial under heterogeneous and changing conditions, despite higher deployment cost, and that applying test-time adaptation after routing further improves performance. The results show that adaptation to unseen distributions can be solved as navigation and composition in a precomputed model space.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.