Empirical-Bayes Truth Tilting for Conformal Inference in Large Language Model Question Answering
Abstract
Large language models (LLMs) are increasingly used as answer engines, yet their tendency to hallucinate makes it important to quantify uncertainty in their answers. Open-ended question answering provides a tractable yet challenging setting for such uncertainty quantification (UQ): sampled responses can be organized into semantic answer clusters, but these clusters are prompt-specific rather than shared across inputs. Each prompt therefore induces a local semantic alphabet of observed answer clusters and an “everything else” atom for unseen answers. Existing query-only conformal methods use missing-mass estimators to assess whether the true answer lies beyond the sampled responses, but they largely equate sampling frequency with truth probability and estimate each instance in isolation. We propose a nonparametric empirical Bayes approach that borrows strength across these apparently unrelated local alphabets. Within each prompt, a Pitman–Yor species model estimates the semantic sampling law; across labeled prompts, an empirical-Bayes truth-tilting layer maps sampling behavior and prompt features to posterior truth mass. The resulting posterior over semantic clusters defines a high-probability nonconformity score, and conformal calibration restores marginal validity. Extensive experiments demonstrate that our Bayesian modeling substantially improves both semantic uncertainty estimation and the downstream conditional coverage of the conformal predictive set for LLM question answering.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.