Morphosemantic Concept Systems: Learning Reusable Discrete Concepts for Compositional Generalization
Abstract
Human languages are characterized by the ability to abstract recurring meanings into reusable concepts and compose them into novel expressions beyond previously observed forms. However, in conventional large language models (LLMs), reusable semantic concepts remain implicit in surface-level tokens and continuous representations, making them difficult to explicitly reuse and systematically recombine in novel contexts. We introduce Morphosemantic Concept Systems (MCS), an intermediate machine language that explicitly represents concepts as discrete identities with a growing inventory, enabling concepts to be reused and recursively composed across different linguistic forms. MCS employs two learnable translators to bridge human and machine languages: an input translator maps natural language into concept-level representations, while an output translator converts generated concepts back into human-readable forms; during training, a gradually annealed teacher provides early contextual supervision to bootstrap concept formation and is removed during inference. Experiments demonstrate that MCS achieves competitive performance with limited training resources. By effectively exploiting and recombining the inherent semantic relationships within sentences through the proposed concept-based representation, MCS validates the potential of resource-efficient semantic abstraction for robust structured reasoning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.