acceptodds
Under review as a conference paper at ICLR 2027

Preference-Based Self-Distillation with Online Contextual Comparisons

Abstract

Privileged context gives a self-teacher information unavailable to its student. The challenge is to turn this information into effective supervision. We introduce Preference-Based Self-Distillation (PBSD), which turns contextual guidance into online comparisons between complete teacher and student responses. A frozen contextual teacher serves two roles: generating the positive response and providing the likelihood reference for scoring both responses. The negative response is sampled from the current student, so the comparisons evolve with its behavior. PBSD optimizes a pairwise logistic objective with labels assigned by generation role. The student receives no privileged context, and training requires no separate preference judge. At each of three Qwen3 scales, PBSD achieves the highest average accuracy across the mathematical benchmarks and the highest tool-call accuracy among the compared methods. On Qwen3-4B, it reaches 64.5% mathematical average, compared with 60.7% for imitation of the same teacher responses, 60.8% with fixed negatives, and 59.3% with an unconditioned reference. These comparisons support the joint use of online responses and contextual reference scoring in the tested setting. Context ablations and pair-correctness diagnostics further characterize the supervision. Together, these results support using contextual self-teachers to construct and score online response comparisons.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.