The Missing Marginal: Marginal Closure for Classifier-Free Guidance in Diffusion Post-Training
Abstract
Classifier-Free Guidance (CFG) combines conditional and unconditional predictions to improve diffusion sampling, yet we observe that it can degrade performance after online reinforcement learning (RL). Examining existing post-training objectives, we find that optimizing either the conditional branch alone or a single fixed-scale CFG output leaves the unconditional branch underdetermined. We term this identification gap the missing-marginal problem. Under shared-target squared-error regression, the unconditional optimum equals the posterior average of the conditional optimum given the noised state, and is determined by the regression target and sampling distribution. We formalize this relation between learned predictions as marginal closure and analyze how closure errors affect CFG. Based on this relation, we introduce Marginal Closure Regularization (MCReg) to train both predictions on a shared target. We derive reward-dependent targets for DiffusionNFT. In on-policy diffusion distillation (OPD), we identify the student's unconditional counterpart and train both branches on conditional teacher targets. Experiments across online RL and multitask OPD show that MCReg improves most evaluated generation metrics under CFG.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.