Latent Stiffness in MeanFlow Training
Abstract
MeanFlow trains one-step generative models using targets that depend on the network’s own time derivative. We interpret this objective as a semi-gradient bootstrap, analogous to temporal-difference (TD) learning: each update fits a target that subsequent updates can change. We introduce a matrix-free Jacobian probe of stiffness, the derivative channel’s parameter sensitivity relative to the output channel. In the TD interpretation, stiffness scaled by the time gap measures how strongly an output-visible parameter change can alter the bootstrapped target relative to the prediction being fitted. The probe therefore exposes target sensitivity that a small prediction residual need not constrain. Across CIFAR-10 and ImageNet-256 checkpoints with up to 677M parameters, stiffness grows by one to two orders of magnitude, especially at large gaps, while the weighted loss remains bounded and FID improves. Controlled experiments distinguish gradual stiffening without bootstrapping from additional amplification under bootstrapped large-gap training whose onset tracks learning-rate warmup. Released α-Flow and Decoupled MeanFlow checkpoints achieve lower large-gap self-consistency error and better one-step FID than MeanFlow despite retaining high stiffness. Increasing large-gap supervision improves the best CIFAR-10 finetuning FID from 6.03 to 3.95 but degrades an ImageNet-256 transformer. These results connect sampling and learning-rate choices to hidden changes in bootstrap sensitivity and give a diagnostic for studying the training dynamics and transfer limits of one-step generators.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.