Evaluating Escalation Signals for Optimal Language Model Routing
Abstract
Cascading a small language model into a large one requires a signal that predicts, before the large model is called, whether the call is needed. We study semantic entropy over sampled answers as that signal, and evaluate it on three tasks of increasing openness: templated arithmetic, GSM8K reasoning, and text-to-SQL. On GSM8K with a Gemma-3 1B/12B pair, semantic entropy predicts the small model's own failures at AUROC , beats a difficulty-calibrated baseline by , and converts into routing gains of up to accuracy over random escalation at matched cost. On synthetically generated chained arithmetic it adds nothing over a regular expression on the question text: a single-family benchmark makes difficulty a free surface feature of the question, so sampling reveals no more than the question already gives away. On text-to-SQL (Gemma-3 4B/12B, ), semantic entropy scores only AUROC pooled, below a free difficulty baseline () – but it is undefined, not weak, on the of items where the samples agree; on the remaining it scores . Splitting on this free-to-compute distinction and calibrating each half before recombining reaches a held-out AUROC of , our strongest result. A retrieval alternative – caching large-model outcomes and querying by embedding similarity – trades signal for cost: it appears to predict large-model success on arithmetic (), but a probe shows this is again difficulty lookup (), and it collapses to chance on GSM8K (), where the same probe returns . It costs always-large against for live sampling at its tuned budget. Framed as prediction rather than diagnosis, semantic entropy earns its keep only where the difficulty of a query cannot be read off the question itself. We report the diagnostics that produced these outcomes – a difficulty baseline, an oracle-headroom check, per-generation cost accounting, a retrieval-difficulty probe, and a regime split – against a measured AUROC noise floor of at these sample sizes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.