Multi-Turn On-Policy Distillation with Prefix Replay
Abstract
Multi-turn on-policy distillation (OPD) gives LLM agents dense teacher supervision on student-generated actions, but requires costly environment rollouts at every update. We propose Replayed-Prefix On-Policy Distillation (ReOPD), which reuses recorded teacher trajectories: the student acts at selected teacher-forced prefixes and receives token-level teacher supervision without interacting with the environment. We identify a prefix trap: student-generated histories improve relevance to the student but can make teacher targets unreliable, while errors compound across turns. A two-sided bound captures the trade-off between student occupancy mismatch and teacher reliability. ReOPD uses a step-decaying sampling schedule that favors earlier, lower-shift prefixes. Across Python-assisted mathematical reasoning and search tasks, with multiple teacher and student scales, ReOPD matches or improves online OPD accuracy, uses zero tool calls during student training, and is at least 4 faster per rollout. Reusing teacher trajectories also supports joint distillation across environments.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.