acceptodds
Under review as a conference paper at ICLR 2027

Implicit Low-Rank Training Dynamics in Directly Parameterized Lipschitz-Bounded deep Networks

Abstract

Recent work has dedicated efforts to developing deep networks with prescribed Lipschitz bounds, in order to limit sensitivity to noisy or adversarial input perturbations. These networks enforce prescribed Lipschitz bounds through built-in structural constraints, typically derived from linear matrix inequalities (LMIs)/ semidefinite programs (SDPs), or imposed through orthogonality constraints on network layers. While these constraints provide certification, they can also make training less efficient. Empirical evidence shows that the training dynamics of large-scale networks often remain in low-dimensional subspaces and that certain initializations can lead to more effective training. Inspired by recent work on emergent low-rank training dynamics of multi-layer perceptrons (MLPs) under gradient descent (GD), we analyze the training dynamics of an existing 1-Lipschitz layer derived from an SDP certification, known as the “Sandwich layer", and show that the training dynamics of its effective weight concentrate within a low-dimensional subspace under GD. We first theoretically characterize this low-dimensional subspace in a shallow network with smooth, slope-restricted nonlinear activations, and empirically demonstrate that similar low-dimensional structure persists in deep networks under more general settings. Building on these observations, we develop a low-rank Sandwich layer that preserves its Lipschitz guarantees. This low-rank parameterization matches the performance of the fully parameterized Sandwich layer across multiple image classification tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.