acceptodds
Under review as a conference paper at ICLR 2027

Prediction Training, Representation Geometry, and Hallucination Detection

Abstract

Prediction training shapes the relative geometry of hidden representations. For an attention model trained by gradient descent, we characterize how representations of contexts associated with the same prediction approach one another. Their remaining differences follow a limiting direction determined by prediction costs, context composition, and the parameter update. This connects prediction training to pairwise representation geometry. We use this component-mixture perspective to develop a pre-generation hallucination detector. For each labeled outcome, the detector reconstructs a test state by reweighting nearby reference states. Its energy penalizes concentrated mixture weights and reconstruction error. It requires no classifier training. The detector achieves a mean AUROC of 79.26% across 21 language and vision-language model–dataset combinations. On visual tasks, it reaches 81.65%, exceeding a supervised pre-generation probe by 2.35 points on the same independent test sets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.