Do Neural Networks Encode Generalizable Patterns in Specific Feature Directions?
Abstract
This paper focuses on the scientific problem of testing whether DNNs encode generalizable patterns in specific feature directions in a verifiable manner. We build on the recent finding that the output score of a neural network can be decomposed into numerical effects of different interaction patterns, some of which can generalize across DNNs and samples, whereas others cannot. We empirically observe that a DNN typically uses only a few feature directions in an intermediate-layer feature space to encode such generalizable interactions. We refer to these directions as reliable feature directions, as they capture generalizable interaction patterns. We identify such reliable feature directions and empirically examine their properties. The identified reliable feature directions usually play a dominant role in classification and are transferable to different samples.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.