Diagnosing How Optimisers Shape Learning through Operator Projections
Abstract
Loss and accuracy track progress, but do not distinguish what an optimiser can change from the direction it chooses or the structure that changes during learning. We introduce an operator-projection framework for bias-free piecewise-linear networks. At each training state, it separates locally attainable operator change, steering within that space, and parameter motion invisible to the sampled operators to first order. Deep linear and MNIST-1D experiments use this reference to examine existing optimisers and a reference-following control. SGD exhibits low alignment despite nearly complete estimated reachability. Increasing reference alignment does not monotonically improve test accuracy, and the control reduces input-dependent gating even when activation slopes remain nonzero. Most strikingly, its pronounced decline in raw operator diversity after fitting nearly disappears after removing the common-logit component that softmax ignores. Across three activations, this component accounts for – of the squared net operator displacement between the measured snapshots. Finite-step checks further distinguish local steering from realised operator change when gates vary. The framework provides a diagnostic instrument for interpreting optimiser behaviour and the evolution of the learned function, rather than a proposed optimisation algorithm.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.