When to Stay, When to Switch: Outcome-Based Evaluation of Adaptive LLM Inference
Abstract
Adaptive LLM inference chooses a model or reasoning budget for each request. Evaluating these choices after an input change requires knowing whether the chosen computation still produces a good answer at an appropriate cost. A route can stay unchanged while its outcome becomes inadequate; a switch can be necessary for a good outcome. We introduce PO-CTRLBench, a paired evaluation framework that measures every available action's quality and cost through repeated execution. These outcome profiles distinguish three requirements: keep an adequate action, allow a change among adequate actions, or adapt when no action remains adequate at both inputs. We show why routes alone cannot identify these requirements, separate outcome failures from selection inconsistency, and give finite-sample certification rules. Our controller, PORC, learns the profiles and selects actions using calibrated conservative utility scores. On two-action reasoning tasks, route consistency and endpoint risk rank five methods in exactly opposite orders. Compared with Outcome-greedy, PORC reduces the primary endpoint-risk sum from 12.1% to 2.2% and improves utility by 0.064 (95% source-cluster confidence interval [0.049,0.079]). Under source shift, changing the decision scores on the same predictor reduces full paired risk from 23.89% to 12.50% and raises utility by 0.040. Paired outcome evaluation thus makes the distinction between stable computation and appropriate computation measurable.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.