A Small-World Perspective on Capability Emergence in Language Models
Abstract
As model scale increases, large language models may exhibit abrupt performance gains on certain tasks—gains that are difficult to predict by extrapolating performance trends observed in smaller models. This phenomenon is commonly referred to as “emergence”. To explain emergence, existing work has investigated it from perspectives such as tasks, evaluation metrics, and model scale; other work has further traced the internal changes within Transformers, for example by tracking the paths of information propagation under a specific input. However, one question remains unanswered: does the Transformer undergo a holistic qualitative change, rather than a single or local change, that gives rise to the dramatic leap in LLM performance? Inspired by the view that “Language Modeling Is Compression”, we propose that model capability should be reflected in whether highly compressed information under a given input can be rapidly reached (finding content seen during pretraining) and rapidly composed (combining multiple related pieces of content). To demonstrate this hypothesis, we borrow the famous six degrees of reachability theorem from small-world theory—namely, that any two nodes in a graph are reachable within six steps, to characterize the internal structural changes of Transformers. Specifically, we first model the Transformer as a directed graph based on the information propagation mechanism of attention; we then extend small-world theory to the directed-graph setting, define metrics that characterize the small-world properties of the Transformer, and describe the quantitative relationship between reachability/composability and these metrics; finally, across different tasks and parameter scales, we demonstrate that changes in the small-world properties of the Transformer's internal organization are highly consistent with the phenomenon of emergence. This is a creative characterization of emergence and may offer inspiration for unlocking greater potential in LLMs. For reproducibility, the code used for the experiments and analyses is available in an anonymous repository: https://anonymous.4open.science/r/iclr2027-A12C/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.