acceptodds
Under review as a conference paper at ICLR 2027

How Multi-Head Transformers Specialize for In-Context Learning: A Theoretical Analysis of Training Dynamics and Generalization

Abstract

In-context learning (ICL) enables transformers to infer task-specific mappings from demonstrations in the prompt. A theoretical understanding of how multiple attention heads learn to specialize during training remains limited. We study the training dynamics and generalization of a multi-head softmax transformer on compositional ICL tasks, where each input consists of multiple latent components that are independently mapped to corresponding output components. We jointly analyze the training of attention and value parameters, and show that their coupled gradient-based dynamics amplify weak initial differences between heads into complementary specialization: each head learns to match and retrieve information according to a different latent component. We establish finite-step convergence and show that the learned model generalizes to unseen combinations of component mappings and to out-of-domain pattern families. Our theoretical insights are validated on synthetic and real experiments.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.