acceptodds
Under review as a conference paper at ICLR 2027

Beyond Hidden States: Component-Aware Confidence Estimation for LLMs

Abstract

Large language models can produce fluent but incorrect answers, with confidence signals that often fail to reflect empirical correctness. Existing estimators rely on output-level signals sensitive to decoding and prompts, or probe hidden states as generic high-dimensional features, underutilizing the functional structure of Transformer components. We propose a component-aware confidence estimator that predicts the correctness of a frozen model's answer from three mechanistic descriptors extracted in a single forward pass: Attention Sink Signature, capturing selected heads' sink behavior; Knowledge Response Alignment, measuring directional agreement between attention and feed-forward residual contributions; and Residual Trajectory Geometry, summarizing residual-stream dynamics through normalized distance and angular statistics. A lightweight branch-attentive probe fuses these descriptors against correctness labels with the backbone frozen. Across five question answering benchmarks and four models from the Qwen2.5 and LLaMA-3.1 families, our method achieves the best AUROC and ECE on most combinations and transfers more reliably than internal-state baselines under cross-domain evaluation. These results show that organizing confidence estimation around component-level structure yields more discriminative and transferable signals than treating internal states as a single high-dimensional feature.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.