acceptodds
Under review as a conference paper at ICLR 2027

Teacher Geometry Shapes Learnability in Teacher-Student Networks

Abstract

Teacher-student systems, in which a teacher neural network generates targets so that a student neural network can learn to implement the same function, are widely used to study learning. However, the structure of the teacher weights is often overlooked by assuming normally-distributed parameters. This hides substantial variations in how learnable different teachers are. We formalize learnability as the success rate of converging to the global minimum and study how it is affected by overparameterization, teacher geometry, initialization distribution, in the context of a bounded data domain. We analytically reduce and systematically characterize the loss landscape to identify two mechanisms through which teacher geometry shapes learnability in small, tractable ReLU systems: (i) In single-neurons systems, the convergence to this minimum is governed by the initial similarity between teacher and student neurons. If the neurons are initially too dissimilar, the student settles into a sub-optimal minimum, which we name out-of-bounds, where the training dynamics push the ReLU activation threshold outside of the bounded data domain. (ii) In two-neuron systems, the alignment between the two teacher neurons determines the existence of suboptimal local minima and shapes their basins of attraction. Motivated by these findings, we empirically study the trainability of larger systems by constructing teacher distributions with different types of alignment between nodes. We discover that increasing alignment between teacher nodes not only reduces learnability but also increases the rate of convergence to out-of-bounds minima. Moreover, convergence rates to these sub-optimal minima can be reduced by increasing the output-layer learning rate, improving learnability. Finally, by training networks on a set of physics regression problems, we confirm that problems that require aligned-nodes solutions tend to have an increased rate of convergence to out-of-bounds minima.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.