READING TRAINING DYNAMICS THROUGH REPRESENTATION DEPENDENCE
Abstract
Neural network training continually reshapes how inputs relate to one another in representation space. Tracking these relationships offers a view of learning that prediction error and gradient magnitude alone do not fully describe. We introduce representation dependence (RD), a training signal that aggregates dependence between hidden representations of different inputs. We analyse what RD captures beyond loss and gradient norm, and how changes in training data affect its evolution. In NanoGPT, adding RD features to a baseline combining loss, gradient norms and representation statistics reduces mean absolute error in predicting validation deterioration 250 updates ahead by 13.6%. RD also tracks representation changes under loss-matched label corruption in ViT and separates instruction-tuning fault windows in Pythia. In grokking, changes in RD accompany delayed generalisation after training examples have already been fitted. These findings show that the evolution of cross-example relationships provides useful information for predicting and interpreting training behaviour.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.