acceptodds
Under review as a conference paper at ICLR 2027

Semantic Uncertainty Probes Across Language Models

Abstract

Semantic entropy is a useful measure of language-model uncertainty, but computing it requires generating and comparing multiple answers. Semantic entropy probes instead predict this uncertainty from a single forward pass, but they are normally trained separately for each model using expensive semantic entropy labels. We study whether a probe trained on one model can be reused on another. Our method learns a lightweight linear map between paired source and target hidden states, without using target-side semantic entropy labels, and then applies the fixed source probe in the aligned target space. Across 21 ordered source–target pairs spanning different model families and scales, transferred probes recover much of the performance of probes trained directly on the target model and sometimes match or outperform them. To understand these results, we partition prediction outcomes according to source-probe errors, transfer-induced changes, and disagreement between source and target uncertainty labels. On average, transfer-induced changes are more often corrective than harmful, while large source–target disagreement limits performance. These findings support the practical reuse of semantic entropy probes and suggest that language models encode partially shared and linearly recoverable uncertainty signals.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.