acceptodds
Under review as a conference paper at ICLR 2027

GOJO: Geometry-preserving Joint Optimization for On-policy Distillation in Diffusion Models

Abstract

In post-training of diffusion models, on-policy distillation (OPD) offers an effective way to combine skills from multiple teachers into a single student. While these teachers are obtained through multiple routes, such as supervised fine-tuning (SFT) or reinforcement learning (RL), a unified regression objective is used to match the velocity. Despite its simplicity, this regression objective can lead to unexpected velocity conflicts among teachers' supervision signals and drift in the student model prior, which causes catastrophic forgetting and teacher matching failure. In this paper, we address this limitation through a local probability-reweighting formulation, where we assign credits to each sampled candidate velocity at student-induced states and update the velocities based on the credits. The credits are constructed through three factors: 1) target teacher agreement that directly guides the student velocity; 2) probability redistribution cost that measures how much the student's geometry is perturbed; and 3) mutual teacher damage that quantifies the impact between teachers. We then integrate these three objectives into a single optimization framework, and introduce GOJO, Geometry-preserving Joint Optimization for multi-teacher on-policy distillation. The resulting optimization problem yields a closed-form solution for the optimal target velocity, which can be used for velocity regression in training. Experiments on SD3.5-Medium and Qwen-Image show that GOJO achieves superior performance in Multi-teacher Distillation with an average 10.5% improvement across tasks on hard-routing OPD and 5.9% on soft-routing OPD, while significantly mitigating the student model prior and avoiding teacher conflicts.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.