acceptodds
Under review as a conference paper at ICLR 2027

RouterShift: Evaluating LLM-Routing Under Non-Stationary Traffic

Abstract

Deployed LLM routers must keep choosing a good model as conditions change: models are constantly updated, the query context drifts, and serving latency rises under load. Standard routing benchmarks assume none of this happens. We present RouterShift, a benchmark that treats routing as an ordered stream with controlled, labeled changes in model quality, availability, query composition, and latency. We investigate twelve scenarios over queries and seven LLMs, across sixteen routing policies and three pool sizes (). Our findings are as follows: i) offline routers lead both the regret and the latency rankings, but not on language and topic shifts; ii) contextual routers, which consume query embeddings, beat their non-contextual counterparts whether online or offline (post-shift regret against at ); iii) topic and language changes are the hardest shifts because they break the learned routing mapping. To speed recovery after a shift, we develop an adaptation layer (AdaptiveShift) that shortens an online router's memory or biases a frozen router's scores when the chosen model's reward drops, removing up to regret per steps. We release construction code, stream specifications, router implementations, and evaluation code that rebuild the corpus from public datasets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.