When Evidence is Missing: Benchmarking Hallucination and Abstention from Incomplete Observations
Abstract
In real-world settings, document-grounded LLM agents must successfully navigate incomplete corpora, whereas most benchmarks assume all evidence is available. We introduce a scenario-agnostic simulation framework that generates document corpora based on a structured latent world. The framework maintains full provenance for all facts, their inference chain, and in which documents they are surfaced, allowing arbitrary subsets of documents to be masked and exactly recomputing which facts remain observable and inferable when certain evidence is missing. We instantiate our framework on four human-validated enterprise scenarios, each producing approximately 1,000 documents. Based on the fact-to-document provenance, we construct question-answering and structured prediction tasks requiring retrieval, multi-document synthesis, logical inference, and abstention when evidence is missing. Evaluating nine LLMs on our tasks using RAG and CLI agent harnesses, we find that models most accurate at answering when evidence is present differ from those that abstain most reliably when it is missing.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.