Causal Core–Periphery Organization and Remodeling in Large Language Models
Abstract
Small subsets of neurons can contribute disproportionately to the behavior of large language models, raising questions about how task-related causal support is organized and remodeled during post-training. We introduce Causal Contribution Profiling, a framework that ranks neurons in the multilayer perceptron (MLP) of Transformer blocks by attribution and measures contribution density by patching activations in progressively larger sets of top-ranked neurons. We study 31 models (0.27B–72B parameters) from five families using a 36-task suite and conduct analyses along three training trajectories. Within the scanned range, we find task-selective core–periphery organization, characterized by a small core with high contribution density and broader support with lower density along the neuron ranking. Core membership is more stable than peripheral membership in 92.0% of 224 eligible task–transition comparisons. Retrospective interventions reveal contributions from future core members before models achieve full accuracy on the evaluated sample groups. These contributions tend to strengthen with behavioral gains. After existing members are excluded, source-stage attribution outperforms layer-matched random baselines in predicting new core and support members in 90.5% and 99.1% of eligible cases, respectively. Together, these findings characterize task support as a contribution-defined organization with differential stability and partially predictable remodeling.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.