Constructive Prediction: The Role of Lexical Code Organization in Language Modeling
Abstract
Character, byte, and subword language models build text from categorical events that are individually valid but may combine into strings with no corresponding entry in a given lexicon. Thus, the prediction interface does not explicitly distinguish such constructions from those that complete a lexical entry. We introduce *constructive prediction*, which instead represents each lexical entry by a short path over a shared, low-cardinality alphabet. In this case, partial paths are intermediate shared states with many possible continuations, and only a completed path can either resolve to a lexical identity or remain unassigned. We instantiate this representation as **Intuitionistic Language Models** (ILM). We first test Flat ILM, which applies a causal Transformer to the coordinate stream. We then extend this design with Full ILM, which adds coordinate-role interfaces and prefix-aware supervision. To assess the effect of constructive prediction, our central experiment asks whether a lexical-to-path assignment constructed from external pretrained embeddings and frozen before downstream language-model training carries predictive structure. To do so, we reassign entries among the same set of paths under the Flat model. Using three assignment seeds, each evaluated with three training seeds, unrestricted reassignment worsens test BPB by on Tiny Shakespeare and on enwik8. All nine paired comparisons favor the original assignment on each corpus. To control for permutation-induced frequency shifts, we use a stricter control that shuffles only entries with equal training frequency, thereby preserving training-weighted coordinate frequencies at each path position. Reassignment still worsens BPB by and , respectively. Beyond assignment, we ask whether the Transformer can exploit path structure. We find that Full ILM has lower mean BPB than Flat ILM at 6.5M and 15.5M on both corpora, and at 100M on enwik8. Full's advantage remains substantial on enwik8 but narrows to BPB at 15.5M on Tiny Shakespeare. Overall, these results show that the organization of lexical identities over shared sequential codes is a modeling variable that can improve prediction under a compact representation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.