Triage: Token-Level Language Model Cascades with Utility-Aware entropy Calibration
Abstract
Large language models (LLMs) have achieved strong performance on complex reasoning tasks, but their growing scale leads to substantial inference costs. Language model cascades mitigate this issue by first using a small language model (SLM) for low-cost generation and selectively invoking a large language model (LLM), aiming to achieve a trade-off between model performance and inference cost. However, existing cascade systems mostly operate at the sequence level and rely on SLM uncertainty as the deferral signal, leading to coarse-grained decisions and a mismatch between uncertainty and deferral utility. In this paper, we propose Triage, a token-level cascade framework with utility-aware entropy calibration (UEC). Triage performs token-level deferral decisions during autoregressive generation and jointly considers the prediction outcomes of the SLM and LLM to calibrate SLM entropy, so that it no longer only reflects model uncertainty, but also better captures token-level deferral utility. Experiments on two model pairs from different families and scales across four reasoning datasets show that Triage achieves a superior cost–performance trade-off over various cascade baselines while preserving the downstream task performance of the SLM.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.