Semantic Entropy Is Estimable Below the Recovery Threshold: Minimax Rates Under a Noisy Judge
Abstract
Semantic entropy scores a language model's uncertainty by the entropy of the distribution from which its sampled answers draw their meanings, and those meanings are never observed: an entailment judge is called on pairs and the answers are grouped by its verdicts. That grouping can fail when the judge is too weak for the available sample size. We model the judge as a channel that, given the meanings, declares each pair equivalent independently with probability when the two answers share a meaning and otherwise. For at most three meanings, a known channel with and a weakly separated judge, , we determine the minimax squared risk of the entropy: , with both constants depending only on . The rate is attained by an estimator that never groups the answers. It reads the collision probability off the agreement density and the triple-collision probability off a centred statistic in the squared degrees, projects the pair onto the set of valid moments, and solves a cubic. The matching lower bound is proved along a family that holds the agreement density fixed while the entropy changes, so every linear statistic of the verdicts has the same mean at its endpoints. With , the entropy is uniformly consistently estimable exactly for , and on the minimax risk is of order . For every , below the recovery threshold , no estimator of the meanings improves asymptotically on assigning each answer to the most likely one, even when given all other answers' meanings. On the entropy is therefore estimable while the grouping is unrecoverable.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.