How Does Preconditioning Guide Feature Learning in Deep Neural Networks?
Abstract
Preconditioning is widely used in machine learning to accelerate convergence on the empirical risk, yet its role on the expected risk remains underexplored. In this work, we investigate how preconditioning affects feature learning and generalization performance. We first show that the input information available to the model is conveyed solely through the Gram matrix defined by the preconditioner’s metric, thereby inducing a controllable spectral bias on feature learning. Concretely, instantiating the preconditioner as the -th power of the input covariance matrix, we prove that, for a single ReLU neuron, increasing monotonically shifts the learned feature's sensitivity toward higher-variance eigenspaces, even though all choices of attain the same limiting population loss. We then empirically investigate how this spectral bias interacts with task-relevant features in three settings: robustness to noise, out-of-distribution generalization, and forward knowledge transfer. Our experiments show that learned representations favor the spectral components emphasized by preconditioning, and that generalization improves when this bias aligns with task-relevant features.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.