acceptodds
Under review as a conference paper at ICLR 2027

FlowCTS: On-policy Continuous Trajectory Supervision of Flow Models

Abstract

While on-policy distillation (OPD) effectively addresses sparse rewards and exposure bias in large language model post-training, its extension to flow models remains underexplored. To this end, we propose Flow Continuous Trajectory Supervision (FlowCTS), which matches subsequent student and reference trajectories initialized from the same student-visited state. Using the integral relation between trajectories and velocity fields, we derive a temporally weighted velocity-matching upper bound and discretize it into practical objectives parameterized by the number of supervision steps. Under a multi-reference setup, FlowCTS-OPD outperforms vanilla KL-based OPD with faster convergence. Compared with the SD3.5-M base model, it improves GenEval from to , OCR from to , and PickScore from to , matching task experts and surpassing the GenEval expert under multi-step supervision, while outperforming mixed-reward RL across all metrics. Further analysis reveals a clear temporal supervision mismatch in vanilla KL-based OPD arising from its auxiliary SDE transition kernels. Beyond on-policy setting, FlowCTS also consistently outperforms vanilla SFT especially on OCR, while increasing supervision steps exhibit a trade-off between richer trajectory information and greater optimization difficulty.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.