acceptodds
Under review as a conference paper at ICLR 2027

View-Decomposed On-Policy Distillation for Robust Reasoning in Large Language Models

Abstract

On-policy distillation reduces training–inference mismatch by supervising large language models (LLMs) on their own generated trajectories. For robust reasoning, this supervision should preserve useful teacher uncertainty while limiting the transfer of prompt-sensitive behavior. We investigate how consistency across task-preserving prompt variations can guide distillation toward more accurate reasoning with less sensitivity to prompt wording. We introduce View-Decomposed On-Policy Distillation (VD-OPD), which compares a frozen teacher's predictions for the same student-generated prefix under two prompt views. Their shared probability mass defines both the learning target and its weight, retaining alternatives supported by both views while attenuating disagreement. VD-OPD uses no answer labels during distillation and adds no inference cost. On six mathematics benchmarks, it raises Qwen3-1.7B-Base's macro-average Avg@8 from 34.8% to 36.7% over EOPD (full), with better unseen-prompt robustness. Gains remain 1.4/1.2 points under two-view/total-FLOPs controls. Cross-view agreement determines VD-OPD's target and weight. Our analysis explains when pooling improves the target and why shared teacher errors can persist.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.