acceptodds
Under review as a conference paper at ICLR 2027

Let's Fail Step by Step: Better Contrastive Distillation with Reliable Worse Teachers

Abstract

Contrastive Distillation, which trains a student policy using a reward defined by the log-probability gap between two teacher policies, has recently gained attention as an effective alternative to On-Policy Distillation. However, prior approaches typically contrast a task-specific, RL-trained teacher with a base policy, limiting scalability to larger teacher models and broader task domains. We instead propose to contrast a strong teacher with a negatively steered version of itself, using a natural-language description encouraging undesirable behaviors. This removes the need for task-specific RL training of the teacher while providing a simple and interpretable way to control contrastive rewards. On mathematical reasoning tasks, we show that steering the negative teacher toward incorrect or non-rigorous reasoning substantially improves the contrastive reward's effectiveness for student training, whereas positively steering the teacher using privileged information or behavioral descriptions yields only marginal gains. We explain this asymmetry through the Steering-Success Rate, defined as the probability that a completion from the steered policy is preferred over one from the base policy. We theoretically show that this quantity directly affects the converged student policy's win rate over its KL-regularized reference policy under reasonable conditions. Through a series of ablations, we demonstrate that the final performance of the student policy closely follows the Steering-Success Rate of its corresponding contrastive reward. In particular, negative steering consistently achieves higher success rates than positive steering, and improving the reliability of negative steering further enhances the resulting reward's training effectiveness. Together, our results establish natural-language steering as a simple and intuitive approach to contrastive distillation, while highlighting Steering-Success Rate as a key factor in designing effective contrastive rewards.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.