acceptodds
Under review as a conference paper at ICLR 2027

Decodable But Not Detachable: Training Data Granularity Determines Parametric Modularity in Large Language Models

Abstract

Do large language models contain domain-specific parametric shells: concentrated, causally necessary neuron populations whose removal selectively degrades a target domain while sparing others? We apply a uniform causal methodology across two domain granularities, three model families (1.5B to 7B parameters), and eight domains. At the academic subject level, zero neurons exceed 60% domain selectivity across 939,008 combined FFN neurons and causal damage matrices are flat, despite domain identity being linearly decodable above 85% accuracy. At the language and modality level, 0.65–1.14% of neurons exceed 60% selectivity, damage matrices are near-perfectly diagonal (ratios up to 595:1), and shell neuron sets are essentially disjoint (IoU < 0.003). Masking code-selective neurons reduces mathematical reasoning accuracy by 16–24 percentage points across all models; chain-of-thought prompting recovers up to 75% of this damage, revealing that the code shell carries both an answer- structuring function and an irreducible reasoning function. Masking Spanish or Chinese neurons leaves mathematical reasoning at or below random. Shell causal strength increases monotonically with model scale, shells are spatially interleaved in a pattern that precludes group-level selective quantization, and amplification of shell neurons degrades performance monotonically, confirming that shells operate at a precisely calibrated functional point. Parametric shells form where and only where training data was modular at the token level.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.