Learning Spherical Polynomials by Learnable Channel Attention
Abstract
We study the problem of learning a spherical polynomial of an unknown degree by training a ReLU network with channel attention by gradient descent (GD) with finite training sample and finite network width guarantee. The target functions are general spherical polynomials with bounded norm and a fixed, unknown degree bound . The training data are i.i.d. uniformly distributed inputs on the unit sphere in with i.i.d. centered noise of finite variance in their responses. The network width is the number of ReLU channels and the size of the diagonal channel-attention matrix. The training process of the network has two stages. In stage one, one ordinary GD step learns channel attention parameters and strictly enlarges the prediction space on the first half of the training data. In stage two, finitely many output GD steps on the second half of the training data achieve expected regression risk as . The feature kernel of the ReLU attention network evolves during training and the learnable channel attention provably reduceds the regression risk, demonstrating the provable benefit of feature learning. With and network width , the sample size and network width are only of logarithmic factors away from their corresponding lower bounds for learning degree- spherical polynomials. Under sub-Gaussian noise, the same algorithm also satisfies a high-probability risk bound with logarithmic confidence dependence. To the best of our knowledge, this is the first guarantee of vanishing regression risk for a finite ReLU network trained by GD with learnable channel attention, uniformly over bounded- spherical polynomials of degree at most , under these sample-size and width budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.