acceptodds
Under review as a conference paper at ICLR 2027

Triage: Token-Level Language Model Cascades with Utility-Aware entropy Calibration

Abstract

Large language models (LLMs) have achieved strong performance on complex reasoning tasks, but their growing scale leads to substantial inference costs. Language model cascades mitigate this issue by first using a small language model (SLM) for low-cost generation and selectively invoking a large language model (LLM), aiming to achieve a trade-off between model performance and inference cost. However, existing cascade systems mostly operate at the sequence level and rely on SLM uncertainty as the deferral signal, leading to coarse-grained decisions and a mismatch between uncertainty and deferral utility. In this paper, we propose Triage, a token-level cascade framework with utility-aware entropy calibration (UEC). Triage performs token-level deferral decisions during autoregressive generation and jointly considers the prediction outcomes of the SLM and LLM to calibrate SLM entropy, so that it no longer only reflects model uncertainty, but also better captures token-level deferral utility. Experiments on two model pairs from different families and scales across four reasoning datasets show that Triage achieves a superior cost–performance trade-off over various cascade baselines while preserving the downstream task performance of the SLM.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.