A Count, Not a Fraction: Hundreds of Examples to Build a Mechanism, Dozens to Trust It
Abstract
How much data does it take to install a mechanism in a language model? Measurements at pretraining scale say the answer is a number of examples rather than a share of the corpus, and they say it without saying why. This paper reproduces that law in a system small enough to write down every path the information can take. A 20M-parameter decoder is trained on a task whose attention mask keeps the query away from the source it must report. A shortcut to a visible copy of that source is open in most training examples, blocked in a chosen number of them, and open but carrying a wrong value in a chosen number more. We expected a relay. What forms instead is a one-hop cache, written into the only positions that both see the source and are read by the query, while the positions after them decode the answer and cannot move it. In a pilot version of the task, formation is governed by the count of examples that force it rather than by their share of the data, and the transition does not move when the dataset grows by an order of magnitude. A model written after the experiments, in which each route’s contribution is a product of a write and a read, reproduces this. An example the shortcut already answers trains both routes in proportion to the attention each receives, so it cannot shift the balance between them, and with gradient clipping the same accounting locates, after the fact, the one delivery schedule that breaks the law, which packs the forcing examples into too few updates. Once the cache exists, a much smaller count of corrupted examples flips which route wins a conflict; across dataset sizes that dose stays a count only when the learning rate decays at the end of training. That flip is a re-weighting applied to every input rather than a detector for the conflicting ones, so most of the answer changes hands while clean accuracy stays where it was.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.