PARALLAX: Automatic Layer Selection For Internal-State Hallucination Detection
Abstract
Large language models hallucinate with confidence, producing text that is fluent, authoritative, and wrong. In medicine and other high-stakes uses, that failure can cause direct harm. Detectors that read a model's own hidden activations promise to catch it before the text reaches a user, and they report high AUROC on standard benchmarks. Those performance numbers are fragile for two reasons. First, in three of the four corpora we study, the detector input already contains the candidate answer being judged, and a sentence-embedding classifier with no access to the model reaches 0.65 to 0.79 AUROC on them. Second, existing detectors are static because they read from a layer selected in advance, often the final layer. That layer is rarely the right one. Cross-validation inside the training data did not select it in any of our 1,000 training folds. The depth that works best also moves from one model to the next, although within a model it barely moves across corpora. PARALLAX searches every layer under cross-validation that never touches the test data, and selects the ones that score best. The effect is large. Across five open-weight models and four corpora, one selected layer beats last-layer probing on every one of those twenty model and corpus pairings, by as much as 0.117 AUROC. It carries to full scale. On the complete corpora, up to 15,090 examples and thirty times the data, it holds on all 100 held-out splits and the largest gain reaches 0.123 AUROC. PARALLAX has the highest AUROC of any detector that reads the model only once, on thirteen of the twenty tests. It keeps its advantage in 58 of 60 cross-corpus transfers, where it was tested on corpora it was not trained on. At a hundred labelled examples and below, it beats the strongest detector that scores tokens individually on every seed. This shows that choosing the layer, with no change to the classifier and no extra forward pass, is enough to put a plain linear probe ahead of every published detector that reads the model once. Which layer a detector reads is therefore a design decision worth making, not a constant fixed in advance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.