acceptodds
Under review as a conference paper at ICLR 2027

Stepwise Semantic Uncertainty in Long-Form Language Generation

Abstract

Semantic uncertainty of large language models (LLMs), i.e., their lack of confidence in the meaning of the generated output, is a well-established indicator of the reliability of short responses. However, in long-form generation, a single response-level estimate cannot distinguish reliable from unreliable units of meaning. Thus, we propose to represent and quantify the semantic uncertainty stepwise, i.e., induced by the next-token distribution at each step during generation, without additional post-hoc response decomposition or access to internal model states that existing approaches rely on. We introduce the concept of an attribute as the topic a unit of meaning refers to, which can be identified as soon as it emerges. By representing the alternative tokens at each step jointly by both their meaning and attribute, we can separate uncertainty due to variation across attributes from uncertainty due to contradictory meanings within attributes. Given that the latter primarily indicates potential factual errors, we leverage our representation to construct an attribute-conditional, kernel-based von Neumann entropy that specifically quantifies this kind of uncertainty. Empirically, we show this is more predictive of token- and unit-wise accuracy in long-form language generation than existing baselines. Since our representation is available before a token is committed, we further utilize it to steer decoding and improve the factual correctness of generated responses.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.