When Answers Stray from Questions: Hallucination Detection via Question-Answer Orthogonal Decomposition
Abstract
Hallucination detection in large language models (LLMs) requires balancing accuracy, efficiency, and robustness to distribution shift. Black-box consistency methods are effective but demand repeated inference; single-pass white-box probes are efficient yet treat answer representations in isolation, often degrading sharply under domain shift. We propose **QAoD** (**Q**uestion-**A**nswer **O**rthogonal **D**ecomposition), a single-pass framework that projects away the question-aligned direction from the answer representation to obtain a question-orthogonal component that suppresses domain-conditioned variation. To identify informative signals, QAoD further selects layers via diversity-penalized Fisher scoring and discriminative neurons via Fisher importance. To address both in-domain detection and cross-domain generalization, we design two complementary probing strategies: pairing the orthogonal component with question context yields a joint probe that maximizes in-domain discriminability, while using the orthogonal component alone supports robust transfer in our cross-domain evaluations. QAoD's joint probe achieves the best in-domain AUROC across all evaluated model-dataset pairs, while the orthogonal-only probe delivers the strongest OOD transfer, surpassing the strongest evaluated white-box baseline for each model by 6.27-14.22 AUROC percentage points on BioASQ at under 2% of generation cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.