acceptodds
Under review as a conference paper at ICLR 2027

Sparse Cross-Layer Modules: A Mesoscopic View of Language Model Organization

Abstract

Interpreting a language model requires connecting individual response patterns to the organization of representations across the network. We study sparse cross-layer modules as a mesoscopic unit for this purpose. Joint sparse coding associates each within-document token response with a decoder direction across the layer stack. In Pythia-410M, the resulting population exhibits depth-dependent differentiation: salient decoder coordinates concentrate in adjacent-layer bands, and response breadth, lexical coherence, and description scores vary across the resulting depth groups. Responses remain heterogeneous within groups, while response similarities extend across modules with little salient-coordinate overlap. Across text domains, the same dictionary exhibits shared usage with distinct preferences, and independently fitted dictionaries retain partial response correspondence. Dictionaries fitted at successive training checkpoints show an early increase in best-match response similarity and a more gradual increase in salient-coordinate overlap. Across three Pythia sizes, matched response subsets retain substantial similarity at comparable standardized-space reconstruction quality. Automated descriptions, activating examples, and selected-module interventions connect these population measurements to inspectable patterns and local output effects. Together, the analyses characterize language-model organization through the distribution of individual responses, their relationships within a population, and their recurrence across inputs and model conditions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.