RouterFactory: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers
Abstract
No single large language model (LLM) is optimal across all queries and budget constraints, making model routing essential for cost-effective LLM deployment. Existing routers span binary quality predictors, cost-aware cascades, graph-based routers, and agentic routers, yet their diverse formalisms and incompatible implementations, coupled with the absence of a standardized evaluation pipeline, hinder fair comparison and further extension. In this paper, we present a unified formulation of LLM routing as a sequential decision process. Under this formulation, a router can be characterized in terms of five types of components: context encoders, model encoders, scoring functions, decision rules, and learning signals. Existing methods can then be organized into three families: single-turn, multi-turn, and personalized routing. Building on this formulation, we develop an automated pipeline that constructs routing supervision by systematically running a pool of candidate models across benchmarks and evaluates routers in terms of both response quality and inference cost under a unified protocol. The resulting benchmark, xRouteBench, spans generic LLM tasks, memory-augmented, vision (image and video), time-series, and personalized routing scenarios. Grounded in the formulation and pipeline, we present RouterFactory, an open-source infrastructure for standardized and modular implementation of LLM routers, where users can add a new router by implementing only a routing method and a loss function and access built-in implementations of 16 representative routers spanning all three families. Using the library and benchmark, we conduct a systematic empirical study of LLM routing and find that learned routers achieve a 17.4% relative improvement over the largest fixed-model baseline, router rankings reverse in favor of lightweight designs under tighter cost constraints, and user-conditioned routing delivers consistent personalization gains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.