acceptodds
Under review as a conference paper at ICLR 2027

Entanglement Is All You Don’t Need

Abstract

This paper primarily focuses on mechanistic investigation, conducting controlled experimentation at a 1B parameter scale to effectively isolate scaling effects, rather than large language model optimization aimed solely at achieving higher benchmark scores. Operationally, we quantify feature entanglement through three measurable properties: (1) decreased intrinsic dimensionality relative to physical dimensionality, (2) increased overlap of activated neurons across different inputs, and (3) lack of hierarchical specialization across layers. The intrinsic dimensionality collapse observed in DeepSeek-V2-Lite-Chat, Gemma, and Qwen indicates premature feature fusion, leading to feature entanglement, which limits the model's representational capacity. To address this problem, we propose a principled architectural design called EPMORE (Explainable Process Mixture-of-Experts) based on the core intuition that the reasoning process is a conditionally dependent feature (from abstract to concrete) orthogonal accumulation process. EPMORE consists of four major components. SPE (Simple Position Encoding) achieves disentanglement between position and content. EPL (Expansive Processing Logic) increases feature capacity by dynamically expanding from extremely low dimensions, making feature disentanglement possible. PIL (Parameter Independence Loss) in the architecture decouples the factor entanglement in weight matrices, thereby achieving feature disentanglement. The MOR (Middle Output Reuse) mechanism allows tokens whose intrinsic dimensionality increases during inference to align with tokenizers at different levels of abstraction; during training, intermediate ground truth labels (not manually annotated) aligned in a hierarchical manner based on the token's level of abstraction (i.e., intrinsic dimensionality) provide intermediate losses that shorten backpropagation paths, accelerating training and reducing computational cost. Ultimately, it enables feature disentanglement between layers. Experiments demonstrate that the EPMORE architecture alleviates feature entanglement, making it possible to provide preliminary evidence toward process-level interpretability. The example of “the capital of France” in the supplementary materials clearly shows how the model derives the result step by step according to the semantic hierarchy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.