Topology and Shared Error Exposure in Large Language Model Lineages
Abstract
As the population of large language models (LLMs) rapidly grows, understanding how shared training sources shape errors across the LLM ecosystem and how developers' training choices create these exposures becomes an important system level question. We develop a lineage representation that maps directed training relationships into error loadings and decomposes model error covariance into three components arising from common tasks, shared base families, and idiosyncratic errors propagated through the lineage. This decomposition shows when concentrated source exposure sustains aggregate error variance as the model population expands. At fixed individual error distributions, different patterns of shared exposure can yield different aggregate errors. We illustrate the mechanism through three canonical lineage topologies. To examine how individual choices shape concentration, we study a one-generation teacher-choice model and derive an explicit condition under which every developer prefers the same teacher, while a diversified teacher assignment reduces aggregate error.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.