acceptodds
Under review as a conference paper at ICLR 2027

Optimize More, Sample Less: Efficient On-Policy Distillation

Abstract

On-policy distillation (OPD) improves a student model by repeatedly sampling from its current policy and distilling feedback from a stronger teacher. However, standard OPD tightly couples sampling and optimization: after each round of rollouts, the student is updated only briefly before new trajectories are collected. This can be inefficient when on-policy sampling and teacher inference are expensive, while optimization on already collected trajectories is comparatively cheap. We revisit this design choice and ask a simple question: should we optimize more before sampling again? We introduce a simple modification to OPD that performs multiple optimization steps on each batch of on-policy trajectories before refreshing the rollout distribution. By increasing the number of optimization steps per rollout batch, the student moves closer to the teacher within each rollout round and extracts more learning signal from every sampled trajectory. On DeepSeek-R1-Distill-Qwen-1.5B across a suite of mathematical reasoning benchmarks, additional optimization reaches a given performance level with substantially fewer rollouts and teacher generations. Compared with vanilla OPD, taking 128 optimization steps per rollout batch achieves a wall-clock speedup and uses % fewer FLOPs, while maintaining or improving final performance. Surprisingly, aggressive reuse of on-policy data remains effective well beyond the conventional one-update-per-rollout regime. Our results suggest that frequent resampling is not necessary for effective on-policy distillation: when sampling is expensive, a simple principle works remarkably well—optimize more, sample less.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.