Compile2Decide: Distilling Frontier Reasoning Models into Fast Calibrated Decision Models
Abstract
Large reasoning models have broad semantic competence but are computationally inefficient for repeated bounded decisions. We ask how much of their judgment can be compiled into a much smaller System One decision model that scores a dynamic candidate set directly, without generating explanations or reasoning traces, and can answer “none of the listed options” when the correct answer is absent. A reasoning-capable LLM teacher (, queried once per input with thinking disabled) scores the full label space plus none; its distribution is projected onto each run-time subset by sending the mass of unlisted labels to none and supplies soft supervision to a ModernBERT-base cross-encoder student, while ground-truth labels anchor the probabilities. The components are standard; the contribution is a decision contract and a pre-registered matched-control measurement. On CLINC150 (confirmatory; BoolQ is descriptive) the student is compared with cross-entropy on ground-truth labels, label smoothing and a permuted-teacher control, and evaluated on calibration, accuracy, latency, throughput, GPU time and energy. Over three seeds, distillation lowered dynamic negative log-likelihood by 0.082, 0.092 and 0.027 nats against the three controls (99% family intervals exclude zero); temperature scaling removed about three quarters of the gap to the ground-truth controls, leaving about 0.02 nats, and accuracy was 0.85 points below label smoothing, so the pre-registered claim that the teacher adds information beyond matched controls is not met. The teacher had no accuracy headroom: the served student was more accurate (0.913 against 0.857), about 4.5 times faster per single request (median 38 against 171 ms) and used a quarter to a third of the teacher's GPU-seconds per decision, breaking even after about 46k–72k requests, an advantage that a label-only fine-tuned student shares; a constrained decoder fine-tuned on labels did not dominate it. What compiled was mainly calibrated probability structure rather than accuracy, which suggests that teacher headroom over labels bounds how much expensive generative intelligence can be converted into reusable, low-latency probabilistic decision functions. The study uses one teacher and one student backbone, and its test sets had been used in an earlier study.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.