acceptodds
Under review as a conference paper at ICLR 2027

Conformal Factuality with Learned Structure

Abstract

As large language models (LLMs) are increasingly used to generate content, evaluating the correctness of their outputs has become essential. We introduce a novel framework for filtering factually incorrect claims from a given piece of AI-generated text while controlling an expected loss, such as the false negative rate (FNR). Our work builds on the recent conformal factuality literature, which leverages conformal risk control (CRC) to select a scalar filtering threshold with distribution-free validity guarantees under exchangeability. However, unlike methods that assess the correctness of each claim in isolation, our approach constructs factuality scores by leveraging the natural language inference (NLI) relationships between all pairs of claims within the text. Specifically, given a piece of text and a raw factuality score function, our algorithm constructs a graphical model over claim correctness, with pairwise factors defined by the estimated NLI scores. Our algorithm then uses belief propagation to approximate the marginal probability that each claim is incorrect, and applies CRC to the resulting updated scores to select a threshold for filtering claims. We demonstrate the effectiveness of our approach on three benchmarks: FABLES, FELM, and a newly annotated Wikipedia-based dataset. The results show that leveraging learned structure between claims substantially improves claim retention over baseline methods while controlling the expected FNR at the nominal level. Remarkably, even when instantiated with open-source models, our method achieves factuality classification accuracy comparable to that of the closed-source baseline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.