acceptodds
Under review as a conference paper at ICLR 2027

Entropy Shortcut in LLM Correctness Probes

Abstract

Large language models (LLMs) achieve strong performance across textual, visual, and other multimodal tasks, yet they frequently hallucinate, producing outputs that are incorrect and unsupported. Preventing such failures is difficult, hence recent works have increasingly focused on detecting them. Two promising approaches that do not modify the model parameters are: (a) sampling the model at high temperature and measuring the entropy of the sampled responses, and (b) training a linear probe on hidden states to estimate the probability of a correct response. Entropy-based detection rests on the hypothesis that sampling entropy and correctness are inversely associated. In this work, we show that a sizable fraction of mispredictions exhibit low sampling entropy. More importantly, standard linear correctness probes fail to distinguish these low-entropy mispredictions from high-entropy correct predictions. We trace this failure to an entropy shortcut- the probes exploit the entropy–correctness association rather than learning a correctness signal that remains stable across entropy regimes. To address this, we propose linear detectors that predict correctness with reduced entropy shortcut. Experiments with four LLMs and five multimodal LLMs across eight datasets, including safety-critical domains, establish the entropy shortcut as a systematic failure mode for LLM probes and demonstrate substantially more robust correctness detection across both entropy-aligned and entropy-misaligned cases.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.