Transformer Provably Hallucinates on Hierachical Classification
Abstract
Large language models (LLMs) can generate fluent and seemingly credible responses that nevertheless contain false statements or unsupported claims, a phenomenon commonly referred to as hallucination. In this paper, we theoretically investigate one particular form of hallucination through a hierarchical sequence classification problem, where a common signal token encodes a coarse label, and a rarer nuance token encodes a fine-grained label. We characterize hallucination as a model having high predicted probability mass on the correct coarse class but near-chance probability on the correct fine-grained label within that class. We prove that a shallow transformer trained by gradient flow exhibits this behavior when nuance tokens are sufficiently rare, and token noise is small. Our results identify a mechanism through which the interaction between data frequency and learned attention can make hallucination a prolonged intermediate phase of training: The model first learns the common signal, but attention weights to the signal tokens further suppress learning from the already rare nuance, producing a long training plateau where the model hallucinates on input sequences containing the fine-grained cue from nuance tokens. Experiments on an LLM-generated synthetic BirdText dataset exhibit analogous coarse-to-fine learning behavior.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.