Route Anything: Quality–Cost Tradeoffs in Model Routing
Abstract
An inference router chooses which model, tool, or combination should handle a task using the current state and records from earlier outcomes. We study three coupled decisions: what capabilities to combine, when to spend more computation, and what information to retain. We use a reference-relative accounting that compares delivered quality with total resource cost, including routing, verification, and information costs, and apply fractional programming to finite-menu routing with a quality floor. On CodeRouterBench and RouterBench, the accounting shows that selection changes the tasks left to a fallback: comparing both portfolios on the same tasks, instead of freezing the current portfolio's conditional outcomes, raises best-candidate selection on CodeRouterBench from 86.3–94.3% to 96.4–97.1%. With a reference policy and an allowed quality drop, maximizing quality per unit cost amounts to choosing a particular balance between quality and cost. The resulting routers remain effective when routing states are inferred from prompts and perform comparably to learned quality–cost routers. With feedback and scoring tests kept disjoint, feedback retries recover 13.9% more failures than fresh retries (0.304 vs. 0.165) across seven models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.