Competence Before Generation, Confidence After: A Two-Stage Self-awareness Framework for Large Language Models
Abstract
While demonstrating remarkable capabilities, Large Language Models (LLMs) often generate fluent yet unreliable answers, making reliability estimation a critical prerequisite for safe deployment. Existing confidence estimators draw on surface-level outputs (e.g., token probabilities and multi-sample consistency) and internal-state probes. We study how internal reliability signals support prospective judgments before generation and retrospective judgments about a generated answer. This perspective echoes human meta-cognition, where monitoring at different stages supports self-awareness. Motivated by this insight, we propose a two-stage self-awareness framework that monitors model reliability from the internal states of a frozen language model. Before generation, the framework estimates the model's competence to answer correctly, defined through sampled-answer correctness. After generation, it assesses the correctness of the particular answer produced. Lightweight heads learn these stage-specific targets with sample-derived soft correctness and consistency supervision, and reuse a single autoregressive generation process at inference. Experiments on seven benchmarks demonstrate improvements in average ranking and calibration over the knowledge-boundary baseline under both in-domain and out-of-domain settings. Across nine in-domain settings, SA improves macro AUROC by 7.51 points; strict out-of-domain evaluation yields 0.7002 ± 0.0064 over twelve settings. Further analyses reveal complementary predictive information in the two stages: pre-generation competence remains informative about correctness after conditioning on post-generation confidence in pooled diagnostics. These findings highlight the value of stage-specific internal monitoring for LLM reliability estimation. Code will be made publicly available to facilitate future research.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.