RSI-Router: Cost-Efficient LLM Routing via Evolution of Subtask-Level Model Assignments with Execution Skills
Abstract
Practical deployment of large language model (LLM) agents requires strong task performance at affordable inference cost. For long-horizon agentic tasks, this performance–cost trade-off can be improved through within-task large–small model collaboration, as smaller models can handle some stages even when they cannot solve the full task. In this paper, we introduce RSI-router, a routing framework that constructs subtask-level model assignments and model-specific execution skills through recursive self-improvement over execution experience. Each iteration consists of four stages: Subtask Mining derives subtask definitions and identification rules from training trajectories; Routing Strategy Evolution proposes and evaluates diverse model assignments; Model-Specific Skill Evolution compares routed and large-model-only trajectories to diagnose failures and develop reusable execution skills; and Pareto-Optimal Router Selection updates the Pareto population from historical and new routers while retaining dominated routers for reuse. Across five agentic benchmarks, RSI-router reduces inference cost by 51.7% on average and up to 82.2% while outperforming the large-model baseline. It also yields a stronger performance–cost Pareto frontier than existing routing baselines. Further analyses show that RSI-router progressively reduces cost over iterations without performance loss, building on previously accumulated skills.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.