Revisiting Momentum Teachers for Pixel-Space Generative Self-Alignment
Abstract
Generative self-alignment methods such as Self-Representation Alignment (SRA) and Self-Flow supervise a generator's intermediate features with targets from a momentum teacher, an exponential-moving-average (EMA) copy of the network. SimSiam showed that a stop-gradient and a predictor can replace this copy for representation learning; we ask whether it is needed for pixel-space image generation. We call the construction without it SiamFlow: both noisy views pass through one network, and the target path, which sees the cleaner view and supplies deeper features, receives no gradient. On ImageNet with JiT-L/16, SiamFlow obtains FID , compared with for the same recipe with an EMA teacher and without alignment. Taking the targets from the current network instead of the EMA copy also lowers FID in the SRA recipe adapted to JiT, and SiamFlow lowers FID at JiT-XL/16 and at . The preferred teacher depends on the recipe: with an MLP + LN projector and neither variance regularization nor token standardization, the online run's FID oscillates and the EMA teacher gives lower FID; adding variance regularization reverses the ranking. Cross-noise instance retrieval ranks the two teachers in the opposite order to FID across training, so this retrieval criterion does not predict the teacher choice for generation. Every model is still sampled with EMA weights.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.