When Data Can Teach: Prerequisite Circuits Gate Learning
Abstract
The same training data can teach a capability to one model and nothing to another. Curricula and data attribution depend on this effect, yet it remains unclear which parts of a model's state determine what data can teach. We test whether a previously learned circuit acts as a prerequisite: holding the training data fixed, we suppress an induction circuit only while a transformer trains on a task that builds on it, and we evaluate every model with the circuit intact. In a two-layer transformer, suppressing the circuit's previous-token head during training blocks learning of long-range copying, and supplying the head's output restores it; without the circuit, the same data teach nothing for 600,000 updates. In pretrained Pythia models from 160M to 6.9B parameters, suppression blocks or delays learning of a retrieval task that builds on induction, whereas suppressing control heads does not. With longer training, most pretrained models learn the task by rebuilding the retrieval circuit from dormant backup heads. Suppressing these heads after training removes the capability, and suppressing them from the start prevents learning at 410M. The delay, 2 to 18 times the unimpaired learning time, depends on how many backup heads remain. Previously learned circuits can thus determine whether, and how quickly, later data teach a new capability.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.