acceptodds
Under review as a conference paper at ICLR 2027

Subspace Learning During Population-Loss Plateaus

Abstract

A nearly constant loss can conceal substantial feature learning. We prove this for two-layer ReLU and leaky-ReLU networks under simultaneous fixed-step population gradient descent from small IID Gaussian initialization. For positive mixtures of damped cubic links in Gaussian Sobolev space, loss stays near one while minimum alignment of the predictor's top- average gradient outer product (AGOP) subspace with the teacher subspace improves by at least . Identically budgeted refitting improves population MSE by more than , and the unchanged trajectory later reduces its own loss. A complementary theorem covers small additive Sobolev perturbations of unequal-weight cubic teachers using projected features. For SwiGLU, at fixed width and dimension as initialization vanishes, we prove leading-AGOP-direction alignment during a loss plateau for square-integrable teachers with a nonzero Hermite component of degree one, two, or three; unrestricted rank-one cubic refit gains hold at a specified width. These guarantees require structured teachers and restrictive parameters. An approximation result further distinguishes subspace recovery from neuron specialization: some interaction targets retain unavoidable prediction error when ridge neurons specialize along shared orthogonal axes within the teacher subspace. Population-moment experiments across 21 teacher links and 50 initializations compare alignment, refit improvement, and trained-loss reduction.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.