Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Space
Abstract
Large language models apply their full computational stack to every token, although the information carried by a sequence varies across spans. We introduce Dynamic Large Concept Models (DLCM), which form variable-length concepts from contextual token representations and process them with a larger concept-level backbone. A lightweight causal decoder uses completed concepts to generate subsequent tokens autoregressively, sharing concept-level computation across multiple predictions. Global compression regularization allows segment lengths to adapt to content while regulating their average. We also introduce a compression-aware scaling law that separates token capacity, concept capacity, and data, together with a decoupled maximal-update parametrization for modules of different widths. A 2.3B-parameter DLCM improves the mean accuracy over a 1.3B dense baseline by 1.01 percentage points across twelve zero-shot benchmarks, with approximately 34% lower estimated inference FLOPs. Both models are trained on one trillion tokens. Gains are strongest on commonsense and question-answering tasks, while compression and token-loss analyses reveal sensitivity to segment granularity. The scaling analysis further predicts lower loss at equal compute through concept-level capacity allocation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.