A Framework for Analyzing DNN Representation Quality: Dynamics of Interaction Generalizability
Abstract
This paper addresses the core challenge in the field of symbolic interaction, i.e., how to define, quantify, and track the dynamics of generalizable and non-generalizable interactions encoded by a DNN throughout the training process. Specifically, this work builds upon the recent theoretical achievement in explainable AI, which proves that the detailed inference patterns of DNNs can be strictly rewritten as a small number of AND-OR interaction patterns. Based on this, we propose an efficient method to quantify the generalizability of each interaction and discover distinct three-phase dynamics of the generalizability of interactions during training. In particular, the early phase of training typically removes noisy and non-generalizable interactions and learns simple and generalizable interactions. The second and the third phases tend to re-learn increasingly complex interactions that are harder to generalize. Experimental results show that the dynamics of eliminating and re-learning non-generalizable interactions are temporally aligned with the dynamics of the training-testing loss gap.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.