Valid but Wrong: Decomposing Language-Model Errors on Closed Vocabularies
Abstract
Practitioners repair language-model output by sampling more, saying more in the prompt, or constraining the decoder. We ask which lever fixes which error, on a task family where the answer is measured exactly: generation against a closed vocabulary, where a machine-checkable oracle separates syntactic error, membership error (hallucinated terms), choice error (wrong real terms), and assembly error (wrong structure). Across more than twenty scored conditions, spanning an intervention ladder on a primary ontology, a cross-ontology validation on a second, and three frontier model tiers, the classes behave in opposite ways. Repeated sampling drives coverage of syntactically valid output to by , while coverage of even half-correct term choice is zero in samples. Constrained decoding raises membership to but degrades structure (paired assembly delta ), and its naive greedy form collapses into repetition loops. A validator repair loop recovers membership () yet lifts choice only to ; a retrieval-only solver peaks at . The strongest condition we find, a frontier model with the vocabulary and per-term descriptions in context, reaches choice , and the same condition on a bijectively renamed vocabulary the model has never seen reaches : familiarity from pretraining is worth (95% CI ), and binding novel symbols to in-context definitions recovers the rest, though replacing meaningful identifiers with arbitrary ones changes more than familiarity alone. Every intervention we test leaves this term-selection residual largely intact; high conformance systematically hides it. In the Linked Open Vocabularies catalogue, 56% of 529 parsed ontologies declare zero refutation-enabling axioms, so for the catalogue's parseable half, even the conformance side of such an oracle is typically unavailable.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.