acceptodds
Under review as a conference paper at ICLR 2027

Transformer Provably Hallucinates on Hierachical Classification

Abstract

Large language models (LLMs) can generate fluent and seemingly credible responses that nevertheless contain false statements or unsupported claims, a phenomenon commonly referred to as hallucination. In this paper, we theoretically investigate one particular form of hallucination through a hierarchical sequence classification problem, where a common signal token encodes a coarse label, and a rarer nuance token encodes a fine-grained label. We characterize hallucination as a model having high predicted probability mass on the correct coarse class but near-chance probability on the correct fine-grained label within that class. We prove that a shallow transformer trained by gradient flow exhibits this behavior when nuance tokens are sufficiently rare, and token noise is small. Our results identify a mechanism through which the interaction between data frequency and learned attention can make hallucination a prolonged intermediate phase of training: The model first learns the common signal, but attention weights to the signal tokens further suppress learning from the already rare nuance, producing a long training plateau where the model hallucinates on input sequences containing the fine-grained cue from nuance tokens. Experiments on an LLM-generated synthetic BirdText dataset exhibit analogous coarse-to-fine learning behavior.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.