Beyond Atomic Claims: Probabilistic Logical Coherence for Long-Form LLM Evaluation
Abstract
Long-form factuality assessment typically follows a decompose-then-verify paradigm: a response is decomposed into atomic claims, each claim is verified against evidence, and the resulting judgments are aggregated into a final score. While effective, this approach evaluates claims in isolation and cannot assess whether they form a coherent argument. We formalize logical coherence as a property of the joint distribution over claim truth values induced by a Markov random field whose factors encode discourse relations expressed within the response. To bridge discourse and inference, we propose a two-level taxonomy that maps twelve PDTB discourse senses onto five inferential couplings. We also introduce LoCoBench, a benchmark that holds claims fixed while varying only their relational structure. Using gold relation graphs, our logical coherence scores satisfy of 202 ordering constraints, compared with – when relations are extracted automatically, identifying relation extraction as the primary bottleneck. Across various baselines, our method is the only one that both orders responses by coherence and remains invariant under meaning-preserving edits. Furthermore, it localizes incoherence through claim-level posteriors rather than just detecting it.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.