What Learning-Rate Susceptibility Measures: Coordinates and Training Routes
Abstract
Learning-rate susceptibility is often interpreted as an intrinsic indicator of training instability, although its value can mix two effects: the coordinate used to perturb the learning rate and changes in the optimization route. We formalize this distinction. Finite predictor-space KL is invariant under coordinate relabeling for a fixed physical intervention, whereas its normalized infinitesimal coefficient rescales with the intervention coordinate. For ReLU–SGD, we give a two-update construction in which bounded smooth-branch Jacobians coexist with an initialization-averaged response because an set of initializations crosses a moving training boundary. We then test the route mechanism in a frozen SVHN/TinyResNet experiment with 20 fresh successor trajectories. Replaying the base trajectory's activation gates and max-pool indices reduces median conditional-quadratic forecast discrepancy in every trajectory by 93.62–99.67%; 15/20 satisfy the predeclared event (one-sided exact ; 95% Clopper–Pearson interval ). Matched controls support route mediation rather than automatic-differentiation dominance, and absolute prediction changes remain small. Thus susceptibility is a coordinate- and protocol-dependent response diagnostic; by itself it does not predict accuracy, divergence, or learning-rate utility.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.