CRANE: Functional Selectivity of Language-Conditioned Neurons in Multilingual Large Language Models
Abstract
Multilingual large language models operate across languages within a shared parameter space, yet it remains unclear how functionally selective language-related internal components actually are. Prior work has mainly identified language-related neurons through activation statistics, but activation selectivity does not directly reveal how these components functionally affect different languages. We propose CRANE, a functional interpretability framework that combines relevance attribution with controlled neuron interventions. CRANE uses AttnLRP to estimate the contribution of MLP components to language-conditioned predictions, separates candidate identification from functional validation, and evaluates target and non-target language behavior under matched neuron-masking budgets to characterize their functional selectivity. We evaluate CRANE on LLaMA2-7B Base across English, Chinese, and Vietnamese using natural language understanding and open-ended generation tasks. Across all evaluated settings, CRANE-identified components exhibit stronger target-language functional effects than those selected by the activation-based LAPE baseline. However, these effects are not strictly confined to the target language, and some non-target languages are also substantially affected. The results indicate that language-related internal components exhibit measurable target-language functional biases, but these biases overlap with cross-lingual functional effects rather than forming fully isolated language-specific modules. We further transfer Base-identified components directly to LLaMA2-7B Chat and observe that part of their functional influence persists after instruction tuning. Overall, CRANE provides an interpretability perspective that moves from correlational identification toward functional selectivity analysis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.