Finding Algorithms in Neural Networks: Convergent Predictions Between Distributed Alignment Search and Speed of Transfer Learning
Abstract
How can we identify the algorithm that a neural network learned on a task? One behavioral method, transfer learning analysis (TLA), is to read it off how quickly it learns related tasks. One mechanistic method, distributed alignment search (DAS), is to read it off from interventions on the network's internal structure. Whether these two methods converge is as of yet untested. We find that, on toy arithmetic and vision tasks, the algorithms identified by DAS and TLA mostly agree: the internal structure of neural networks predicts how quickly they will transfer to a related task (and conversely, speed of transfer predicts internal structure). Crucially, only DAS with non-linear localization agrees with TLA in our vision task. This means that TLA isn't merely a description of the network's behavior. DAS and TLA are two convergent sources of evidence for identifying the algorithm implemented by a neural network.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.