acceptodds
Under review as a conference paper at ICLR 2027

Not All Privilege Is Distillable: Calibrating Teacher Corrections for Few-Step Video Generation

Abstract

Privileged distillation uses clean video continuations to supervise a few-step generator on student-visited denoising states. We find that the largest evaluated evidence budget yields lower mean generation scores and motion activity than an intermediate budget, even when teacher and student share the first frame. To examine reference dependence, we introduce Privilege-Induced Teacher Variance (PITV), an offline measure of normalized variation in teacher corrections across screened references at fixed teacher weights and student states. Higher PITV is associated with poorer full-privilege transfer. For training, Mode-Calibrated Privileged Distillation (MCPD) tests support within a single continuation. It compares corrections from a progressively enlarged evidence subset and the full selected budget, using directional agreement and normalized displacement to weight the full correction at each token. Poorly supported targets move toward the teacher prediction conditioned only on the prompt and first frame. On 4-step Wan2.2 and 8-step LongCat-Video, MCPD improves VBench-I2V Total Score over timestep-dependent scalar shrinkage by 0.6–0.9 points and also outperforms adapted DFD. Controls on Wan2.2 show that matching residual energy or mixing the same teacher queries with fixed weights does not recover MCPD's performance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.