Gram–Schmidt Breaks the Gradient Structure of Riemannian Gradient Descent: First-Order Retractions on the Stiefel Manifold
Abstract
Orthogonality constraints are commonly found in deep learning, appearing in applications such as unitary recurrent networks, orthogonal convolutions, and Stiefel-constrained layers. These constraints are typically enforced by retracting each update using a thin QR factorization, which involves Gram-Schmidt orthogonalization. However, QR factorization is known to be only a first-order retraction. We demonstrate that this has structural implications for the optimizer itself. Specifically, for Riemannian gradient descent on the Stiefel manifold, we derive, in closed form, the tangential defect introduced by QR factorization at second order in the step size. Notably, this defect disappears when using the polar retraction. Furthermore, we prove that using explicit objectives validated through exact rational arithmetic, the -form dual to this defect is not closed. As a result, gradient descent employing QR factorization does not support any modified loss at first order in the step size, in contrast to methods based on the exponential map and other second-order retractions. Since the defect vanishes at critical points, it goes unnoticed in standard local convergence and stability analyses. Additionally, because Gram-Schmidt processes columns sequentially, the optimization based on QR is not equivariant under permutations of the weight columns. This means that in a Stiefel-constrained convolutional network and an orthogonal recurrent network, permuted copies of the same model diverge within the first few updates, while updates based on polar retraction remain consistent to machine precision. We cannot determine whether this lack of equivariance affects final accuracy since, over the course of full training, even the polar control eventually loses equivariance due to accumulated rounding errors. Our primary concern is with well-posedness and reproducibility. Lastly, we show that the polar retraction can be computed by a short Newton–Schulz iteration built entirely from matrix multiplications. This method restores equivariance at the tested sizes and performs faster than thin QR factorization at larger dimensions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.