The Factorization Was the Circuit: Chinese Remainder Theorem in Transformers
Abstract
Emergent algorithms can arise in neural networks trained on simple tasks. Twenty small transformers (1-layer, no LayerNorm) trained to solve over a range of primes exhibit perfect generalization after grokking. We reverse engineer the emergent mechanism and find that the model first rewrites the answer as , and , where and is a primitive root of . It then splits that product using the (CRT): computing on independent streams, where are the prime-power factors of ; then recombining the residues in the unembedding. Similar to prior work, we learn that the model uses a sparse set of "key" frequencies, which index the irreducible representations. Unlike prior work, this set is completely deterministic, as it is the union of disjoint bands of the Fourier spectrum at the multiples of . We call these bands the CRT bands. Through causal tests, we show that each band controls its own corresponding stream independently. Furthermore, for prime-power streams (e.g., ), the irreducible representations of these groups are filtered by order, dual to the quotient tower: . With the a priori knowledge of these irreducible representations and their frequencies, we are able to show the streams' independence causally, as well as isolate and observe the mechanism's circuitry formation during grokking.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.