Vanishing Initialization Turns Diagonal Tensor Gradient Flow into Greedy Matching
Abstract
We study gradient-flow (GF) dynamics of canonical polyadic (CP) tensor factorization from vanishing initialization. For order- models with and targets diagonal in mode-wise orthonormal bases, we prove that the limiting training trajectory is governed by greedy bipartite matching. Each atom–sector pair acquires a random activation clock determined by its microscopic initialization; processing these clocks in increasing order, subject to atom and sector exclusivity, determines the successive learning events. Away from activation times, the tensor output converges to the corresponding staircase trajectory. Complex modular addition is an exact equal-sector instance after Fourier diagonalization, so the activation clocks are i.i.d. Combining this reduction with known Random Edge asymptotics yields an explicit distribution-free large-system loss curve in the edge-exposure parametrization; we further derive its asymptotic conversion to physical training time. For real weights, conjugate Fourier modes instead form two-dimensional sectors. We characterize their rank and border-rank cascade and conjecture that the full GF dynamics is organized by the corresponding interleaved cascades.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.