acceptodds
Under review as a conference paper at ICLR 2027

What Should Black-Box LLM Confidence Predict? Estimating the Correctness of a Fixed Answer from Multiple Signals

Abstract

Black-box confidence methods score an answer by asking the model, resampling it, perturbing the input, or consulting other models, but need not score the same object. We freeze each question–answer pair before measurement and ask what survives new models and questions. Signals derived from the target itself, self-report and resampling, mostly measure commitment, including commitment to stable errors. Other models supply external corroboration. Under simultaneous model and question holdout, a small head that fuses self-report with peer evidence beats retrained text encoders and a released correctness model, and a frozen text encoder given the peer measurements nearly recovers the fusion. Asking the same peers to judge the frozen answer reads more of that evidence than counting their agreement on five of six datasets. A monotone environment link transfers the ranking while the probability scale stays local; peer value is governed by the panel's separation between right and wrong target answers; and the correctness of a unanimous consensus leaves the unlabeled agreement law unchanged. Separation predicts that stale agreement ages when an upgrade erodes it. Three frontier models showed it, the second time in a preregistered replication: the same previous-generation peers weakened whether counted or asked to judge, refreshed peers restored both readouts, above self-report at a moderate gap and indistinguishable from it at a wide one, and pooled stale peers diluted it. Freeze the answer, read capable peers as judges, refresh them when the target moves, audit the scale locally, and never equate unanimity with truth.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.