Gradient Descent Without Repeated Backpropagation via Lazy Low-Rank Jacobians
Abstract
Backpropagation is the standard and highly effective mechanism for computing gradient-descent updates in neural networks. We explore whether part of its repeated derivative computation can be amortized by exploiting structure that emerges in large, overparameterized models. Our approach builds on two such properties: the *lazy regime*, where the Jacobian changes only mildly during training, and substantial low-rank structure in the Jacobian. Together, these suggest that derivative information can be compressed once and reused across optimization. We develop practical tools for testing these properties without explicitly forming or storing the full Jacobian or Hessian, including a Hessian–Jacobian measure of laziness and a scalable rank-estimation procedure based on a small subset of model parameters. We further show that, under suitable parameterizations, parameter-efficient fine-tuning inherits the laziness of the base model. Building on these observations, we construct a training procedure that reuses a compact low-rank Jacobian representation to approximate standard gradient-descent updates without repeated network-wide backpropagation. The forward computation remains unchanged, while the costly network-wide backward pass is replaced by inexpensive updates over a compact Jacobian representation, making parameter updates nearly negligible compared with standard backpropagation. Our goal is not to replace backpropagation as a general-purpose training method, but to show that gradient-based learning can be reorganized around the mathematical structure of overparameterized networks, substantially reducing the need for repeated backpropagation. Experiments show that this alternative computational organization remains viable in practice on realistic architectures and datasets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.