acceptodds
Under review as a conference paper at ICLR 2027

LLM Uncertainty Revealed in Behavioral Boundary Effects

Abstract

Large language models may produce fluent and confident answers even when those answers are unreliable. Existing black-box uncertainty estimators typically rely on repeated sampling from the same model or disagreement across an unordered collection of models. The former may fail when the model consistently reproduces the same error, while the latter may introduce unnecessary disagreement and inference cost due to differences in model architecture and training distribution. In this work, we investigate whether an ordered family of small models can serve as a more effective uncertainty probe, and propose **AnchorAUC**. Given an input, a target model first generates an anchor answer, after which a sequence of progressively smaller probe models independently answers the same input. Their semantic retention relative to the anchor forms a directional degradation trajectory along an ordered model-scale axis. AnchorAUC integrates this trajectory into a reliability score: answers that remain stable across a broader range of probe levels receive lower uncertainty, whereas answers that drift earlier exhibit a behavioral boundary associated with higher uncertainty. We evaluate AnchorAUC on five question-answering benchmarks, four target models, and two probe families. AnchorAUC achieves competitive AUROC and AUARC across diverse model–task configurations while requiring substantially lower inference cost than multi-anchor sampling. Further analyses of probe order, correctness transitions, and probe granularity show that the ordered trajectory contains information beyond generic cross-model disagreement.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.