Implicit Language Representation of LLMs: A Formal Language Learning Perspective
Abstract
When a large language model (LLM) is trained to learn a language, is it learning the underlying language representation (henceforth, also structure) precisely, or is it also becoming proficient in other related languages? This question offers insight into the LLM’s implicit world model. To answer this question precisely, we need a learning environment for LLMs where we can define the language and its underlying structure. We propose probabilistic formal language as the learning environment, which provides greater control than natural language. Next, we need a proficiency test that considers the possibility of generating new strings both inside and outside the language. Our answer is a local discriminative test that operates on individual strings: a string is learned if the LLM generates it with higher probability than every neighboring out-of-language string, and language proficiency is the fraction of strings learned this way. Empirically, within a language, an LLM learns strings better if they are structurally closer to the training data. Beyond individual strings, we characterize the LLM’s implicit structure, i.e., its formal grammar here, by the languages it generalizes to beyond the target. We find that the LLM generalizes better to out-of-distribution languages that relax global structure than to ones that relax local structure. Simultaneously, the implicit grammar is close to grammar rules governing local structure, but deviates from rules governing global structure, a pattern we also confirm empirically in natural language, where it learns local word structure better than global sentence- or paragraph-level structure.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.