acceptodds
Under review as a conference paper at ICLR 2027

A Local Theory of Gradient-Based Rank-One Recovery in Nonlinear Low-Rank Models

Abstract

Sequential rank-one methods build a low-rank update one component at a time. In the linear case, the residual, which is the part of the target update not yet recovered, directly reveals the next component. Nonlinearity hides that residual, but we can still compute the loss gradient from training data. Can this gradient reveal a useful direction? Motivated by sequential rank-one LoRA, we study this question through a direction-locked recovery procedure in a single-layer nonlinear teacher-student model. This provides a nonlinear counterpart to classical Hotelling deflation for memory-bounded fine-tuning and tasks with unknown target rank. At each stage , the procedure extracts a rank-one direction from the gradient of the squared-error adaptation loss , fits a scalar amplitude , and freezes the update before proceeding. For a Gaussian Leaky-ReLU teacher-student model with rank- target and slope , we prove that the rescaled negative population gradient equals the current residual plus a nonlinearity-induced bias controlled by a sector angle. The empirical gradient on a fresh batch concentrates around this surrogate at rate . Under sector control and gap-relative bounds on nonlinear bias, sampling error, and accumulated drift, this yields local per-stage recovery. After stages, the conditional multi-stage bound controls the final parameter error by the oracle rank- tail plus accumulated local errors. Sharpness examples show the worst-case need for spectral separation in top-dyad recovery and covariance control in the frozen-feature bridge, while a complementary gap-free bound gives sufficient conditions for one-stage parameter-residual decrease. Our formal contribution is a local residual-surrogate theorem for this direction-locked procedure. Empirically, we test the surrogate mechanism, observe residual progress outside the top-dyad sufficient regime, and illustrate the tradeoff between rank and active optimizer-state memory on frozen image features.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.