acceptodds
Under review as a conference paper at ICLR 2027

A Theoretical Analysis on the Emergence of Implicit Curriculum in Transformers

Abstract

Empirical studies of language-model pretraining suggest an implicit curriculum: localization and retrieval capabilities often emerge before capabilities that compose retrieved information. Existing theory, however, typically analyzes these ingredients separately, studying attention-based token selection with simple readouts or nonlinear composition on fixed representations, leaving open whether joint training can itself induce a structured acquisition order. We study this question in a stylized shared-attention multi-task Transformer, in which a softmax attention layer, shared across tasks, feeds task-specific ReLU heads. The model is trained by stochastic gradient descent on a fixed mixture of two tasks whose input sequences contain informative tokens among distractors: a token-selection task, labeled by the sign of a single informative token, and an XOR task, labeled by the XOR of the signs of two informative tokens. Both tasks are present throughout training, and the shared attention-weight parameter and both heads are updated in parallel, without an externally imposed curriculum. Our main result proves that the joint trajectory nevertheless exhibits a definite learning order: token selection reaches small test error while the XOR error is still bounded away from zero, and the XOR error becomes small only at a later iteration. The mechanism is that token-selection learning amplifies attention to a positional signal shared by both tasks, while parity cancellation and cumulative stability control keep XOR updates from disrupting this amplification. Experiments on synthetic data from the model and MNIST-based token sequences support the predicted learning order. As supporting empirical context, preliminary measurements on open intermediate large language model (LLM) checkpoints show a compatible pattern in which perception-like skills emerge earlier than composition-like skills.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.