acceptodds
Under review as a conference paper at ICLR 2027

Prime by Prime: A Reference for Emergent Modularity

Abstract

Neural networks can break a task into subtasks and learn to solve each independently, in a different part of the model. But what determines whether this modular organization emerges during learning remains poorly understood. One obstacle is that the relevant subtasks and their internal representations are often unknown. We address this by studying transformers trained to compute and , where is a product of distinct primes. These tasks admit a simple decomposition: represent the inputs through their remainders and , perform arithmetic separately for each prime, and reconstruct the unique answer modulo consistent with the results, as guaranteed by the Chinese remainder theorem. We ask whether transformers discover this decomposition through training. We find that models trained jointly on addition and multiplication (i.e., learning the ring ) consistently do so, with evidence at three levels: (i) errors occur approximately independently across primes, (ii) each prime’s computation occupies its own representational subspace, and (iii) these computations largely rely on distinct groups of neurons. This modular organization critically depends on the training task: addition alone produces entangled or partially modular solutions that vary across seeds, including fused subtasks and asymmetric dependencies between primes. Why training on the ring induces this modular solution, while addition alone does not, remains an open theoretical question and an interesting direction for future work.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.