Past-Self Guidance: Internalizing Temporal Contrast in Flow Models
Abstract
Training produces a history of predictors, but generation usually uses only its final model. At sampling time, adding the component of a current-past prediction difference perpendicular to the current prediction improves samples but not regression loss. We introduce Past-Self Guidance (PSG), which adds the full prediction difference to the flow-matching target. Both predictions use the same noisy input, one from the model's recent exponential moving average and one from an earlier ordinary-training checkpoint. Sampling needs only the trained model. On ImageNet 256 x 256, one setting improves FID-50k by 34-55% across four SiT scales at 400k updates, without classifier-free guidance (CFG). At XL/2, FID falls from 16.89 to 7.70 at similar recall. PSG also has lower FID than ordinary training that uses about twice as much compute. Applied only in the last 10% of a 2M-update run, PSG reaches CFG-free FID-50k 5.73, against 10.10 for ordinary training. It needs that run's earlier checkpoints and about 6% more training compute. Activated at 1M updates, it reaches FID-50k 1.99 with CFG at 2M. It combines with REPA and improves diffusion and unconditional generation. Ablations and sampling interventions show that the gain comes mainly from the direction of this difference and disappears when the difference is taken from another sample.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.