Entropy Shortcut in LLM Correctness Probes
Abstract
Large language models (LLMs) achieve strong performance across textual, visual, and other multimodal tasks, yet they frequently hallucinate, producing outputs that are incorrect and unsupported. Preventing such failures is difficult, hence recent works have increasingly focused on detecting them. Two promising approaches that do not modify the model parameters are: (a) sampling the model at high temperature and measuring the entropy of the sampled responses, and (b) training a linear probe on hidden states to estimate the probability of a correct response. Entropy-based detection rests on the hypothesis that sampling entropy and correctness are inversely associated. In this work, we show that a sizable fraction of mispredictions exhibit low sampling entropy. More importantly, standard linear correctness probes fail to distinguish these low-entropy mispredictions from high-entropy correct predictions. We trace this failure to an entropy shortcut- the probes exploit the entropy–correctness association rather than learning a correctness signal that remains stable across entropy regimes. To address this, we propose linear detectors that predict correctness with reduced entropy shortcut. Experiments with four LLMs and five multimodal LLMs across eight datasets, including safety-critical domains, establish the entropy shortcut as a systematic failure mode for LLM probes and demonstrate substantially more robust correctness detection across both entropy-aligned and entropy-misaligned cases.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.