Smaller Models, Better Rejects: Preference Distillation Scaling
Abstract
Preference distillation commonly treats a teacher response as preferred and the student's own response as rejected. This practice rests on two assumptions: the student's own failures are the most informative negatives, and rejects must come from a model as large as the student, which makes reject generation increasingly costly as students scale. We find that neither holds: across students from 7B to 72B, smaller frozen models generate rejects with less inference compute yet train stronger students than the student's own rejects, both before and after sequence-level knowledge distillation, on code generation and mathematical reasoning. To analyze this finding, we ask which reject distribution most improves a given student and derive a finite-horizon utility bound in a linearized feature model of DPO. The bound characterizes a favorable region of reject distributions and motivates three interventions. First, because net transfer in the bound is linear in source mixtures, we randomly mix rejects from a smaller model and from the student-scale model; performance rises as the smaller model's share grows. Second, because the bound predicts that task content independent of the prompt survives when rejects are detached from their prompts, we reassign rejects to other prompts and further shuffle their code tokens; both still outperform length-matched gibberish, so part of the gain comes from task structure rather than from errors specific to each prompt. Third, the analysis shows that selecting candidates with lower likelihood under the SeqKD reference helps when higher-likelihood candidates carry less useful contrast, as may happen for near misses that share features with correct responses. We reselect each source's candidate rejects by this likelihood; lower-likelihood selections outperform higher-likelihood ones for every source. Together, these results suggest a simple design principle: effective rejects preserve task structure while limiting coupling to the reference policy, and smaller frozen models provide both at low cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.