Cost–Accuracy and Budget Adherence for Out-of-Distribution Reasoning Routing
Abstract
Reasoning routers decide when a language model should spend additional computation on a reasoning response. A useful router must both allocate reasoning to requests where it improves accuracy and continue meeting a compute budget when the request distribution changes. Cost–accuracy curves measure the first requirement, but not the second: deployment uses a threshold selected before the shift. We evaluate both requirements across two prediction models, diverse reasoning tasks, and unseen-dataset shifts by selecting thresholds on source validation data and applying them unchanged at test time. Scaling supervised routers from 0.8B to 27B yields only marginal in-distribution cost–accuracy gains, no consistent out-of-distribution gains, and greater variability in budget transfer. Moreover, routers with nearly identical cost–accuracy can differ substantially in realized cost and in how much their routing rates change, showing that ranking quality alone is insufficient. Existing supervised routers use only the question as input, whereas heuristic routers use signals from the direct prediction. We combine these approaches by also giving supervised routers the completed direct prediction and its confidence. This significantly improves in-distribution and out-of-distribution cost–accuracy at both 0.8B and 27B, and substantially reduces out-of-distribution budget-adherence error at 27B. Across our evaluated settings, the 0.8B router with these inputs remains the practical default. Reasoning routers should therefore be evaluated using cost–accuracy, cost adherence, and routing-rate adherence together under distribution shift.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.