Normalized Perceptual Pullbacks for One-Step Generative Modeling
Abstract
One step image generation places most distribution correction into training, making the construction of regression targets central to optimization. Feature guided drifting provides a natural way to build such targets, but the effect of transferring a feature space correction back to a pixel generator remains unclear. We show that, under a matched frozen field and scaling, feature target regression and an unnormalized pixel pullback induce the same local first order generator update. The pulled back correction is unevenly scaled across samples: with DINOv3-L, the largest 10% of samples account for 41.8% of the total correction magnitude. We therefore normalize each pulled back direction before forming stop gradient pixel targets. Normalizing in feature space leaves substantial variation after the encoder Jacobian, while normalization after the pullback directly controls pixel correction scale. Experiments show the same pattern across representation families, encoder capacities, image resolutions, and generator scales. On ImageNet, the method reaches FID 1.47, 2.04, and 3.18 at 256, 512, and 1024 resolution. On BLIP3o 60k, it reaches cFreD 3.03 and CLIPScore 23.23 under one step inference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.