When Does Loss Normalization Preserve Stationarity in Consistency Learning?
Abstract
Consistency models are commonly trained with pseudo-Huber losses, whose gradients normalize each sampled correction before averaging. Normalization can move a model away from its intended flow map even when the expected unnormalized gradient vanishes there. We characterize when it preserves this stationary state. For a fixed joint law with integrable Jacobian-weighted corrections, stationarity holds at every positive floor, or smoothing constant, exactly when the Jacobian-weighted correction has zero mean conditional on the correction magnitude. This radial-balance condition requires cancellation within each group of equal correction magnitude, so it is stronger than cancellation of the overall mean. A smooth Gaussian mixture violates it at the exact map, and one of three networks trained to approximate this map meets prespecified criteria for accuracy and for separating the normalized field from the linear one. For fixed features and quadratic error, we derive sufficient conditions under which one step reduces the error, while a step of the same size from the same state under the reversed correction law increases it; on 16 fresh feature seeds, every resolved forecast agrees with the computed outcome. Across six paired tuning seeds from one pretrained diffusion model, two-step Fr\'echet inception distance favors zero-floor normalization over a squared-loss control with matched update lengths; three earlier continuation studies give mixed results that depend on the update policy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.