How Multi-Head Transformers Specialize for In-Context Learning: A Theoretical Analysis of Training Dynamics and Generalization
Abstract
In-context learning (ICL) enables transformers to infer task-specific mappings from demonstrations in the prompt. A theoretical understanding of how multiple attention heads learn to specialize during training remains limited. We study the training dynamics and generalization of a multi-head softmax transformer on compositional ICL tasks, where each input consists of multiple latent components that are independently mapped to corresponding output components. We jointly analyze the training of attention and value parameters, and show that their coupled gradient-based dynamics amplify weak initial differences between heads into complementary specialization: each head learns to match and retrieve information according to a different latent component. We establish finite-step convergence and show that the learned model generalizes to unseen combinations of component mappings and to out-of-domain pattern families. Our theoretical insights are validated on synthetic and real experiments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.