Training as a flow of distributions
Abstract
Analyses of neural-network training usually examine one aspect at a time, the dynamics of the optimizer, the energy of the loss landscape, or the geometry of the underlying space (representational, functional, or parameter), each in isolation. To unify all in a single apparatus, we use the Wasserstein gradient flow, the steepest descent of a free energy in Wasserstein geometry, which coincides with the Fokker-Planck equation and is realized in discrete time by the Jordan-Kinderlehrer-Otto (JKO) proximal scheme. This scheme unifies the energy, geometry, and learning dynamics in one object, making their interaction measurable. The measurements that this framework relies on and employs are high-dimensional and lack a ground truth to validate them, and they might deviate from the correct measurement without any indication of failure. We therefore validate each measurement on small, analytically tractable models, in which the answer is known in closed form. Employing generative classifiers versus discriminative models, and focusing on measurements introduced and known in the related literature, such as calibration, we compare these models and validate these results: (i) a predictor's calibration is controlled by the transport distance of its parameter distribution to the posterior, tracked along the flow down to a finite-sample floor; (ii) the geometry of the landscape sets the rate at which calibration is reached, governed by the slowest mode, (iii) the crossover of discriminate over generative classifiers occurs as a result of misspecification and in the case of the correct assumption, the crossover does not happen.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.