acceptodds
Under review as a conference paper at ICLR 2027

Learning When to Signal Uncertainty in LLMs

Abstract

Large language models (LLMs) often produce confident yet incorrect answers, which can lead to risky failures in real-world applications. We study whether post-training can make a model's self-assessment explicit: when the model is uncertain, can it be trained to signal so within its own response? A central design question is where in the response this signal should be exposed: during reasoning, while the answer is still being formed, or at the end, once the answer has been produced. We study both. For end-of-reasoning self-assessment, we train the model to verbalize a confidence score for its response, with the aim of high confidence on correct answers and low confidence on incorrect ones. For during-reasoning self-assessment, we train the model to emit the marker `<uncertain>` as a warning, using final-answer correctness to supervise whether it appears. Across factual question-answering tasks, confidence training improves calibration, and generated warnings identify answers more likely to be incorrect. Both reports support selective retrieval to revise completed answers. We also examine internal representations through confidence-related geometry, confidence-digit probabilities across layers, and probes across model families. On shared responses, these analyses characterize associations between internal states, expressed confidence, and answer correctness.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.