CoRoute: Learning Routing Frontiers as Ordered Policy Hulls
Abstract
A routing service exposes several operating points at once, and callers expect them to agree: a query accepted at a lower tier should stay accepted when coverage or compute allowances increase. Independent per-allowance optimization, however, offers no such guarantee: a more generous tier can withdraw a query that a stricter tier accepted. We treat the routing frontier as the object to learn, rather than one operating point, and require the learned family to preserve acceptance across tiers. For a shared score and a finite library of threshold-action policies, CoRoute learns tier policies within an ordered policy hull, using first-order stochastic dominance to couple their threshold distributions. One shared draw nests realized acceptance for every input while leaving model selection flexible. The order condition is linear, enabling globally optimal joint fitting within the library. Its fitted-risk cost is zero exactly when independently optimal policies admit a coherent selection. As a reusable policy layer, CoRoute keeps router predictors and LLMs fixed and adds only tens of microseconds once scores are available. Independent service data can additionally certify coverage and bounded resource budgets. Experiments on published LLM cascades and multiple routing architectures establish withdrawal and evaluate the component. In an independent external evaluation, CoRoute reduces the mean fraction of requests losing service on at least one upgrade from about 8% to zero at a mean accuracy cost of 0.04 percentage points, while reducing the largest tier losses relative to a strong single-cut alternative.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.