Infrasampling: A framework for language modeling without hallucinations
Abstract
In this paper, we investigate efficient, general, and scalable methods for selective prediction in open-domain language modeling. We reconceptualize hallucinations as *data points whose probability under the model is too high*, and distinguish *epistemic* and *aleatoric* hallucinations by analogy to uncertainty quantification. From these definitions follows a general framework for trading off between epistemic hallucinations and abstentions in language models, which we call *infrasampling*. The key idea is that we sample from a *lower bound* on the distribution—rather than an average density estimate—and treat the slack as abstention. We instantiate this framework for autoregressive language models in the soft distillation setting, using new output heads that capture the model's epistemic uncertainty about the teacher distribution. The resulting hallucination-abstention frontiers dominate strong heuristic baselines and suggest compounding gains with scale.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.