Beyond Sampled Answers: Conformal Prediction for Language Models via Latent-Space Expansion
Abstract
Conformal prediction (CP) provides prediction sets with finite-sample marginal coverage at level and has recently been applied to language models (LMs). For open-ended generation, sampling-based CP scores calibration examples using the empirical frequency of their ground-truth answers across repeated generations. However, if repeated generation misses the ground-truth answer for more than an fraction of calibration inputs, these methods may fail to attain the target coverage or return the uninformative entire answer space. To recover ground-truth answers missed during generation, we introduce a hybrid nonconformity score that combines empirical generation frequencies with discrepancies measured in an encoder-induced latent space. The hybrid prediction set therefore extends beyond sampled answers through a latent region that permits checking whether an answer is included but does not enumerate the answers it covers. We enumerate the additional answers covered by this region within a finite vocabulary to obtain an explicit semantic prediction set. To reduce set size, we partition the discrepancy space into cells and prioritize those that recover more missed ground-truth answers while admitting fewer incorrect vocabulary answers. Under exchangeability, the hybrid prediction set guarantees marginal coverage of at least . At a target coverage of , its empirical coverage averages 90.25% across five datasets versus 39.59% for sampling-based CP. The semantic prediction set retains this guarantee when the vocabulary contains all missed ground-truth answers.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.