acceptodds
Under review as a conference paper at ICLR 2027

When LLM Confidence Fails: External Calibration for Tabular Predictions

Abstract

Large language models (LLMs) are increasingly applied to structured clinical predictions, making reliable uncertainty quantification essential for safe deployment. Evaluating across three model families (both open and closed weights), three prompting conditions and nine binary clinical prediction tasks with over 16,000 combined patient records, we demonstrate that verbalized confidence from LLMs is fundamentally uninformative regarding instance-level correctness on tabular data. By formally disentangling two commonly conflated evaluation metrics, we show that an LLM carrying zero correctness information can deceptively attain a classifier AUROC of 0.806, while its true self-knowledge inverts to below random chance on complex etiologies like acute kidney injury. To address this, we introduce a scalable, two-call external calibrator that predicts per-instance correctness using task-level descriptors, rank-aggregation statistics, and the divergence between the LLM's self-reported feature attribution and a data-derived reference. Because our calibrator requires no model internals, it is directly applicable to proprietary APIs. While cross-task transfer without target adaptation reveals that external calibration fails on certain targets due to concept shift, adapting the calibrator with a small few-shot target set resolves this degradation. Crucially, source-task pretraining reduces the required target labeling budget by roughly an order of magnitude, achieving correctness AUROC gains (up to +0.363) and providing a robust, lightweight guardrail for LLM deployment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.