Implicit Low-Rank Training Dynamics in Directly Parameterized Lipschitz-Bounded deep Networks
Abstract
Recent work has dedicated efforts to developing deep networks with prescribed Lipschitz bounds, in order to limit sensitivity to noisy or adversarial input perturbations. These networks enforce prescribed Lipschitz bounds through built-in structural constraints, typically derived from linear matrix inequalities (LMIs)/ semidefinite programs (SDPs), or imposed through orthogonality constraints on network layers. While these constraints provide certification, they can also make training less efficient. Empirical evidence shows that the training dynamics of large-scale networks often remain in low-dimensional subspaces and that certain initializations can lead to more effective training. Inspired by recent work on emergent low-rank training dynamics of multi-layer perceptrons (MLPs) under gradient descent (GD), we analyze the training dynamics of an existing 1-Lipschitz layer derived from an SDP certification, known as the “Sandwich layer", and show that the training dynamics of its effective weight concentrate within a low-dimensional subspace under GD. We first theoretically characterize this low-dimensional subspace in a shallow network with smooth, slope-restricted nonlinear activations, and empirically demonstrate that similar low-dimensional structure persists in deep networks under more general settings. Building on these observations, we develop a low-rank Sandwich layer that preserves its Lipschitz guarantees. This low-rank parameterization matches the performance of the fully parameterized Sandwich layer across multiple image classification tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.