Dynamic LLM Routing: When to Admit a New Model?
Abstract
To leverage newly released Large Language Models (LLMs) cost-effectively, dynamic routing methods onboard new LLMs using observations of whether they answer the queries correctly (i.e., correctness observations). Such observations may mislead the routers to redirect queries from existing LLMs that would answer correctly to a new LLM that answers incorrectly. Consequently, this worsens the overall system accuracy. However, existing dynamic routing methods focus solely on router updating, leaving the deployment decision: *whether to admit a candidate LLM*, unaddressed. To bridge this gap, we propose the *Dual-Evidence Admission Filter (-Filter)*, an admission decision method that can operate on top of existing dynamic routing methods. -Filter combines *pool-history evidence* with *self-simulated evidence*. The former quantifies the reliability of predicted routing gains using residuals from pseudo-onboarding trials; the latter uses the candidate LLM's observations to assess whether redirecting queries to the candidate improves accuracy. These dual-channel evidences are then jointly used to determine whether to admit the candidate. Experiments across multiple benchmarks and routing methods demonstrate that -Filter reduces harmful model admissions while retaining beneficial routing updates.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.