acceptodds
Under review as a conference paper at ICLR 2027

SE-OPSD: Self-Elicited On-Policy Self-Distillation For Open-Ended Task

Abstract

On-policy distillation provides dense token-level supervision on model-generated trajectories. Recent on-policy self-distillation (OPSD) removes the need for external teachers by supervising a model with a privileged copy of itself. In verifiable domains such as mathematics and code, this privilege can be supplied by reference answers or verified reasoning traces. In open-ended QA, however, response quality depends not only on factual content but also on recognizing missing context, calibrating uncertainty, respecting task-specific boundaries, and communicating appropriately. A reference answer may leave these response-framing requirements implicit. A model may identify them when explicitly prompted to analyze a query, yet fail to apply them consistently when answering directly. These elicited judgments constitute natural privileged information: a teacher that has already formed them can supervise a student that has not. We propose SE-OPSD, a reference-free self-distillation method that uses query-specific response context as training-time privileged information. For each query, a frozen teacher copy first elicits a compact, structured checklist of response requirements through complementary views, forming Self-Elicited Privileged Context (SEPC). The student observes only the original query, while the teacher conditions on both the query and SEPC to provide token-level supervision on the student's on-policy trajectory. SE-OPSD thus uses additional inference computation to construct its own supervision and distills these response-framing judgments into single-call answer generation. On HealthBench Consensus, SE-OPSD improves Qwen3-1.7B from 67.86 to 72.88 and Qwen3-4B from 82.15 to 86.58, outperforming reference-answer SFT and reference-conditioned OPSD baselines. Gains in context awareness, communication quality, and completeness, together with improvements on general-domain WildBench, support self-elicited response context as an effective source of supervision for open-ended QA. Our code is available at https://anonymous.4open.science/r/SE-OPSD.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.