CALIBER: Calibrated System-One Routing for Adaptive System-Two Reasoning
Abstract
Reasoning models improve performance on difficult tasks but incur substantial latency and computational cost when invoked indiscriminately. We study whether a fast calibrated decision model can learn when deliberative reasoning is actually necessary. CALIBER introduces a System-One/System-Two architecture in which a lightweight decision model observes the question and one cheap System One pass and selects among a direct System One answer from a 4B model, a direct answer from a 27B model and extended reasoning at either scale, and can defer when no action is likely to succeed. The router returns a calibrated probability of success for every action rather than an unconstrained textual plan, enabling explicit risk-sensitive policies that trade expected success against cost. We train the router on the realized outcomes of executing every route on every question, and compare it with confidence cascades, a prompted language-model router, conventional classifiers, a binary win classifier, random routing and an outcome oracle. Experiments on MMLU-Pro and SuperGPQA measure task success, calibration, risk–coverage trade-offs, routing regret and GPU cost, both in distribution and under held-out-subject and cross-benchmark shift. The central hypothesis is that calibrated decision routing can preserve most of the capability of expensive reasoning systems while substantially reducing their invocation rate. The hypothesis holds only in part. At half the cost of always reasoning with the 27B model, retains 83% of that route's accuracy gain over the System One answer, short of the pre-specified 90% target, while invoking System Two on 41% of questions. In the pre-specified comparison under the same budget cap it is 2.9 points more accurate than the best confidence cascade in distribution (81.5% versus 78.6%; 95% interval [1.3, 4.5] points), and about 2 points at matched realized cost. The gain comes from per-action success heads over a text encoder for the large direct and large reasoning actions, not from temperature scaling, System One confidence features or the small reasoning action, which lowers accuracy under this policy. Under held-out-subject and cross-benchmark shift, reasoning becomes longer and operating points transferred from the source calibration split miss their budgets; refitting temperatures and operating points on a small target calibration set restores budget adherence, with no significant difference from the cascade.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.