Meaning Before Words
Abstract
Predicting words and preserving meaning are different tasks. We study this gap in a finite language model where many strings express one meaning. The probability chain rule can already represent any latent semantic generator; the obstruction lies in the decisions and resource allocation built around it. For every fixed temperature below one, we construct a full-support binary model whose semantic error tends to zero, while both token-level and sequence-level temperature sampling have error tending to one. Even exact sequence maximization fails. We give a sharp finite-length bound for the local effect. Under a shared memory limit, an optimal next-token predictor can discard an available semantic bit; we compute the exact cost of protecting it. A sharp log-loss bound separates semantic guarantees from length-normalization artifacts. Finally, a meaning-first interface makes decisions independent of verbal form when its decoder is faithful. The results separate semantic decision-making from its autoregressive language realization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.