Lifting Cheaper Models: Self-Evolving Routing for Cost-Efficient Agent Services
Abstract
As agent services built on large language models (LLMs) move into production, their cost accrues with every request, much of it the premium for top models. Model routing lowers this premium by serving a request with a cheaper model whenever one matches the model the user asked for, but it is confined to the models already in the pool. When every cheaper model fails a recurring task, the router can only pay for the expensive one or accept worse results, and its savings stop at the capability boundary of the cheaper models. We observe that these failures are rarely random: a cheaper model tends to break at the same step across tasks of the same kind, and the execution history already records how a stronger model gets past it. We introduce RouteSage, a self-evolving router that edits its own routing table through capability patches, validated entries that let a cheaper action serve a family of tasks on behalf of the requested model. A direct patch reroutes the family to a cheaper model that already suffices. A guided patch pairs a cheaper model that falls short with an Experience, a reusable procedure distilled from past executions, creating an action the pool never had without retraining any model. Direct patches recover savings that existing models allow, and guided patches lift the models that fall short. Across web interaction, tool use, and code generation, RouteSage exceeds requested-model quality in every domain while lowering total online cost, including the cost of building and validating every patch, and on MiniWoB++ it cuts online cost by 29.6% and raises success by 3.4 points. Served from a frozen state, it delivers the highest quality of any method in every setting and cuts cost by 57.4% when GPT-5.6-sol is requested on MiniWoB++, whereas no baseline saves more than 5.2% without losing quality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.