acceptodds
Under review as a conference paper at ICLR 2027

Reward-Guided Knowledge Distillation

Abstract

The goal of knowledge distillation is to improve a student model’s performance using output samples of a teacher model. The performance can be measured, yet this explicit signal has not been used to control the impact of individual teacher samples during training. In this work, we establish Reward-Guided Knowledge Distillation (RG-KD), a unified and principled framework that weights the distillation loss according to the expected downstream utility of the teacher's predictive distribution. RG-KD further extends to sequential prediction through trajectory-level reward weighting and critic-based step-level value estimation. Extensive evaluations across five diverse tasks demonstrate that RG-KD consistently outperforms unweighted and entropy-weighted baselines as well as recent adaptive and selective distillation methods by effectively weighting teacher signals and aligning the student with high-reward outcomes. RG-KD also improves existing approaches through plug-and-play integration.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.