acceptodds
Under review as a conference paper at ICLR 2027

Predictable but Not Actionable: A Decision-Gap Analysis of Static LLM Routing

Abstract

Large language model (LLM) routing routes queries to balance quality and cost, yet static routers plateau below achievable quality. Why this plateau persists, and how routing should be evaluated, remains open. Existing evaluations are flawed because they score routers as predictors, not policies whose switches determine value. A corrected standard separates genuine gains from measurement artifacts before deployment. We formalize the decision gap D(q), the oracle improvement over a stated prior, and decompose value into switch rate and per-switch gain. On RouterBench (36,497 queries), we evaluate query-only readers, prefill probes, and a two-stage cascade. Our protocol separates oracle-aware from selection-valid evaluation via deployable inputs. Shuffle controls and nested cross-validation govern comparisons. Ten-seed sweeps, calibration audits, and leakage-free replay validate the findings. We find headroom on 18.1% of queries, readers flipping a third of decisions for near-zero gain, and one cascade gaining under oracle-aware but losing under leakage-free replay.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.