DualTrace: Probing Evidence Integration for RAG Hallucination Detection
Abstract
Retrieval-augmented language models can rely on retrieved text yet still misrepresent it: context dependence alone does not establish faithfulness. We introduce DualTrace, a supervised detector built around two complementary views of evidence integration: how removing context changes a fixed response computation, and whether predictions decoded from hidden states agree across depth. DualTrace replays the same response with and without context, measuring output changes, hidden-state movement, and intermediate-to-final prediction disagreement, then pools these signals for response-level classification. On RAGTruth and HalluRAG, DualTrace improves on the evaluated final-layer controls and transfers across datasets and models. Averaging feature-family detector scores also outperforms selecting a single family, supporting their complementary predictive value. Across settings, hallucinated responses exhibit lower early uncertainty but greater late-layer disagreement. Context sensitivity further correlates with fixed-response log-probability changes under activation replacement. Together, these results connect a practical detection representation with measurable changes in evidence-sensitive computation and prediction agreement across depth.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.