Language Models Cannot Tell When a Definition Is Missing
Abstract
Contracts, technical documentation, and online communities give familiar words meanings of their own. When the defining text is missing from a language model's input, the model answers by the word's usual meaning, and nothing in its output signals the error. We study these silent misreadings on four corpora with diagnostic questions whose answers follow from a term's local definition. Frontier models answer 13% to 30% of these questions wrongly without the definition. The errors concentrate on familiar words given a new meaning, often repeat across ten samples, and persist across model generations, as 70% to 94% of the newest model's errors are also made by an older one. Sampling agreement and stated confidence rank them with AUROC 0.46 to 0.56 on redefined common words, and the calibrated confidence of a decision model reaches only 0.64 to 0.66 until the definition is supplied. Checking a small open model against the answer implied by the source ranks frontier models' errors with AUROC 0.72 to 0.85, where term frequency and rarity reach at most 0.63, and places two thirds of a contract library's estimated errors in a quarter of its questions. The source definition repairs most errors, whereas the model's own gloss of the term does not. Models can follow a definition they are given but cannot recognize when one is missing, so local definitions have to be located and supplied from the source.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.