acceptodds
Under review as a conference paper at ICLR 2027

How Much Correctness Signal Is in the Top- Log-Probability Channel? An Attribution Study of Single-Pass LLM Uncertainty

Abstract

Reliable LLM deployment needs uncertainty estimates that are both faithful and cheap, yet the strongest methods resample the model, while single-pass methods trade accuracy for cost. We ask how much response-correctness signal a single generation's token-wise predictive distributions actually carry once three everyday factors are accounted for: response length, supervision, and model access. To answer, we build PULSE, a supervised sequence probe that reads only the per-step top- log-probabilities of a single generation, without access to hidden states or weights, and compare it with reference probes that read progressively more of a response, across five LLMs and three mathematical reasoning benchmarks. Most of the black-box signal comes for free: response length alone is a strong predictor, and a zero-shot entropy score comes within about two AUROC points of PULSE. Supervision adds nothing over these free scalars on its own; reading the full top- channel adds about a point, and modeling sequence order a smaller increment on top. White-box hidden states, by contrast, open a much larger gap. Within the black-box setting, PULSE is the strongest option we test, never significantly worse than the best baseline and matching a five-pass perturbation-based estimator at one-fifth of its cost. The resulting attribution shows where black-box uncertainty signal comes from, and how little of it lies beyond free, zero-shot scores.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.