acceptodds
Under review as a conference paper at ICLR 2027

Unpacking the Compute Efficiency of Strong-to-Weak OPD against GRPO

Abstract

On-policy distillation (OPD) trains a student on its own rollouts under a frozen teacher's per-token supervision. Does this dense signal reduce the training FLOPs needed to reach a target accuracy compared with reinforcement learning? We systematically compare vanilla reverse-KL OPD with Group Relative Policy Optimization (GRPO) across task difficulty, student and teacher size, domain, and optimization choices, measuring held-out accuracy against training FLOPs. With an external post-trained teacher, OPD generally loses to GRPO in accuracy per training FLOP, even when the teacher is further trained with GRPO, and often suffers response-length inflation. Yet a different teacher changes the picture: in our math setting, OPD recovers the accuracy of the student's own GRPO checkpoint with 50× fewer distillation FLOPs than direct GRPO needs to reach that accuracy. On a single domain, however, the GRPO checkpoint is already the desired model, so distilling it adds to total compute. Across domains, though, separate GRPO checkpoints hold different capabilities, and consolidating them into a single model has a distinct advantage. Multi-teacher OPD consolidates per-domain GRPO experts into a single student that matches each expert's accuracy, with the distillation stage costing only 5.2% (on 2 domains) and 8.0% (on 4 domains) of the specialist training FLOPs. Counting specialist training and distillation together, this pipeline reaches joint GRPO's performance on the same data mix at 42.0% fewer total FLOPs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.