acceptodds
Under review as a conference paper at ICLR 2027

Pay for Rollouts Only When Needed: Compute-Aware Routing for Language Model Distillation

Abstract

On-policy distillation addresses exposure bias by training a student language model on its own generated responses, with supervision from a teacher. However, generating and scoring these responses, specially with frontier models, is more expensive compared to off-policy distillation using cached teacher responses. We investigate whether selectively allocating on-policy supervision can reduce distillation cost while maintaining student quality. In samples where a student favors one teacher-supported choice while ignoring others (i.e., anchored collapse), off-policy training with cached teacher probabilities can help recover those alternatives. In contrast, when a student assigns substantial probability to choices the teacher considers unlikely (i.e., support escape), the student's trajectory may not match the cached prefixes in off-policy distillation, making fresh rollouts via on-policy distillation helpful. We formalize this in our proposed compute-aware routing (CAR) approach for distillation, choosing among three actions: skipping an update, cached off-policy and on-policy distillation. CAR compares student and teacher probabilities on cached prefixes to predict action utilities that account for compute costs. In particular, we compute directional support statistics from cached prefixes to distinguish anchored collapse, for which off-policy may suffice, from support escape, for which fresh on-policy supervision may provide additional value. We also provide a routing-regret bound conditional on utility-estimation accuracy and an idealized support-corridor bound linking escape probability to loss-surrogate discrepancy under a common trajectory distribution. In a one-seed GSM8K study, CAR achieved 62.40% accuracy versus 62.24% for vanilla on-policy distillation while routing 71.02% of samples to fresh on-policy supervision, indicative of cost savings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.