The Token before the Value is the Key: How Hybrid Architectures Organize Induction Circuits
Abstract
Hybrid language models combine efficient and global layers, but how these layers develop complementary roles in induction remains unclear. We analyze induction in terms of three functions: carrying predecessor information, matching a historical source by content, and copying its value. Using layer-type-agnostic paired probes at a common block-update interface, we track Carrying and Matching in recurrent–global and local–global hybrids. In both settings, Carrying concentrates in efficient layers, while Matching develops in global receivers. The measured local contribution is concentrated at lag one, on the token immediately preceding a historical value. Interventions that alter predecessor support–lag-one masking, convolution removal, and reduced early learning rates–shift the distribution of Carrying and Matching across stages. Source-key restoration and fixed-value selection further trace the receiver's dependence on the prepared source. Even with the architecture and training corpus held fixed, early learning-rate interventions produce different natural-text prediction and recall outcomes. Varying local windows and using induction-enriched training text also changes when functional Carrying and Matching emerge. Together, these results suggest that learning to carry predecessor information helps determine where and when Matching develops, linking architectural design and training evidence to induction circuits and recall. Code is available in https://anonymous.4open.science/r/bind-match-copy-F4C8/README.md.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.