acceptodds
Under review as a conference paper at ICLR 2027

Calibrating Verbalized Confidence at Any Output Token Position

Abstract

Large language models often report confidence that does not match their accuracy, and existing methods typically calibrate it at the before-answer (BA) or after-answer (AA) position. This leaves open what confidence denotes at an arbitrary output position and how it should be calibrated. We define the target at token as the conditional probability of eventual answer correctness given the current prefix; these values form a Doob martingale whose endpoints are BA and AA confidence. We then show empirically that coupling a BA report with answer generation can alter subsequent accuracy, so such reports need not estimate the solver's success probability under a plain reasoning prompt. Position-Agnostic Confidence Elicitation (PACE) therefore decouples the two: it samples and grades solutions once, then trains a confidence readout at randomly selected prefixes with a Brier reward, which we prove is uniquely maximized by the conditional correctness probability for a fixed solver and position-selection rule. We further analyze how group normalization in group relative policy optimization (GRPO) alters this objective. Mean subtraction preserves the expected reward-gradient direction, whereas standardization yields a comparison semi-gradient that, with two samples, induces threshold-type best responses; we characterize larger groups through structural identities and ablations. Across five backbones and seven test sets, PACE matches or outperforms strong baselines designed for one or two positions. Finally, confidence-based routing demonstrates the value of position flexibility: deferring at BA saves more cost than routing on AA confidence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.