Same Fixed Point, Different Learning Dynamics: Error Transport in Finite-Step Flow Learning
Abstract
Training objectives for few-step generation can share a solution yet differ in the error feedback they supply during learning. For a frozen teacher, integrated endpoint labels stay fixed, whereas stopped differential labels return a directional derivative of the student's error. We show that an order- horizon error receives diagonal residual gain , with material transport coupling different orders. These unequal gains let one global step size contract some modes and destabilize others in idealized target-space learning. In image networks, the kernel and AdamW reshape these gradient gains before they reach the learned function. Paired training across three image domains favors endpoint supervision; at five sampling steps, CIFAR-10 FID is 7.590 versus 21.133 for stopped-JVP training after 3,000 updates. Endpoint's advantage persists at equal measured time, including label production on shared, reusable training states. Choosing a target thus also chooses how different horizon orders are corrected.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.