A PHYSICAL THEORY OF BACKPROPAGATION: EXACT GRADIENTS FROM THE LEAST-ACTION PRINCIPLE
Abstract
Backpropagation (BP) computes gradients through successive forward and back- ward passes. We derive the same gradients from an extension of Hamilton’s action principle that describes dissipative systems using two coupled states. Their mean represents network activations, while their difference carries the sensitivities needed to compute gradients. The task loss supplies the output error, and the two states evolve together through interactions between neighboring computational stages. These interactions use the same transposed local derivatives as BP. At equilibrium the gradient is exactly BP; updating all stages together with unit-step Euler reaches this equilibrium within 2Lupdates for a feedforward chain of Lstages. The formu- lation also lets us change how the gradient is reached. Grouping operations into larger blocks reduces the number of update rounds. Adding inertia and choosing its settings from the system’s decay rates speeds up the tested recurrent gradient computations while preserving their equilibrium. The formulation opens the door to using tools from physics to study and design learning dynamics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.