Beyond Linear Concepts: Discovering and Aligning Non-Linear Concept Manifolds in Large Language Models
Abstract
Understanding information processing in large language models (LLMs) requires dissecting the geometric organization of their internal token representations. While existing mechanistic interpretability (MI) methods seek to extract human-understandable concepts, they are constrained by a strong linearity assumption—an assumption increasingly challenged by evidence of non-linear feature manifolds. In this work, we move beyond linear concepts by adapting Non-Linear Multi-Dimensional Concept Discovery (NLMCD) from computer vision to token-level LLM activations, modeling concepts as low-dimensional manifolds. To enable quantitative comparison of these independently discovered concept manifolds across layers and models, we introduce a concept-based alignment (CBA) score, a generalized Rand index that measures geometric proximity without requiring explicit feature matching. Our analysis yields six key findings: (i) a neighboring-layer sanity check shows CBA is more sensitive than PCA- or CKA-based linear baselines; (ii) layer×layer alignment matrices reveal two characteristic block structures concentrated in intermediate and late layers, consistent across models and obscured by linear metrics; (iii) concept composition remains syntax-dominated through most of the network before giving way to increasingly mixed syntactic–semantic concepts in later layers, with increasing output-orientation toward the final layers; (iv) multilingual concept sharing between English and Mandarin is training-dependent rather than universal — strongest in Qwen, weaker in Llama, and absent in GPT-2; (v) inter-model alignment mirrors this structure: strong correspondence between same-family Qwen models of different scale, consistent with the block structures in (ii), but weak alignment across model families; and (vi) tracing representational alignment across the training stages of the Tulu-3 pipeline shows alignment is highest between adjacent stages, validating CBA, with the largest shift occurring between the base model and the SFT stage, while subsequent preference-alignment stages (DPO, RLVR) leave early layers largely unchanged and RLVR mostly preserves DPO's concepts in late layers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.