acceptodds
Under review as a conference paper at ICLR 2027

CollabMask: Identifying Neuron Collaboration for Interpretable LLM Adaptation

Abstract

The internal interpretability of large language models (LLMs) remains a critical challenge, particularly for understanding how distributed neurons coordinate across tasks. In this work, we identify neuron collaboration, an empirical phenomenon wherein specific neurons form closely connected neighborhoods that are systematically co-activated. We validate this collaborative pattern through two complementary lines of evidence: semantic alignment in independently obtained functional interpretations of collaborating neurons, and selective co-activation on targeted tasks and token types. To apply this structure to adaptation, we propose CollabMask, which generates collaboration-aware gradient masks to induce structured sparsity during continual adaptation. Experimental results show that leveraging this structure improves standard adaptation on mathematical benchmarks by an average of 2.7%. Furthermore, in continual learning scenarios, preserving these collaborative subnetworks helps mitigate catastrophic forgetting, surpassing existing baselines on new tasks while retaining previously acquired knowledge. By connecting neuron collaboration, structured sparsity, and task-specific dynamics, our findings provide an interpretable perspective for understanding LLM adaptation and advance the empirical explainability of multi-task learning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.