LLM Uncertainty Revealed in Behavioral Boundary Effects
Abstract
Large language models may produce fluent and confident answers even when those answers are unreliable. Existing black-box uncertainty estimators typically rely on repeated sampling from the same model or disagreement across an unordered collection of models. The former may fail when the model consistently reproduces the same error, while the latter may introduce unnecessary disagreement and inference cost due to differences in model architecture and training distribution. In this work, we investigate whether an ordered family of small models can serve as a more effective uncertainty probe, and propose **AnchorAUC**. Given an input, a target model first generates an anchor answer, after which a sequence of progressively smaller probe models independently answers the same input. Their semantic retention relative to the anchor forms a directional degradation trajectory along an ordered model-scale axis. AnchorAUC integrates this trajectory into a reliability score: answers that remain stable across a broader range of probe levels receive lower uncertainty, whereas answers that drift earlier exhibit a behavioral boundary associated with higher uncertainty. We evaluate AnchorAUC on five question-answering benchmarks, four target models, and two probe families. AnchorAUC achieves competitive AUROC and AUARC across diverse model–task configurations while requiring substantially lower inference cost than multi-anchor sampling. Further analyses of probe order, correctness transitions, and probe granularity show that the ordered trajectory contains information beyond generic cross-model disagreement.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.