acceptodds
Under review as a conference paper at ICLR 2027

EnterpriseSaga: A Billion-Token Benchmark for Evidence Retrieval and Reasoning over Evolving Enterprise Knowledge

Abstract

Enterprise question answering requires connecting evidence across systems while distinguishing activities that share participants, terminology, and business context. We introduce EnterpriseSaga, a billion-token benchmark for evidence retrieval and reasoning built around a synthetic B2B SaaS company. It contains 966 questions spanning 18 months and 12 enterprise systems, with 1.213 billion tokens in the exported serialization. Our Coarse-to-Fine Enterprise Simulation Framework develops projects and recurring operations within a shared company history and generates artifacts grounded in their activities. To represent the repetition and overlap of real company operations in finer detail, question-conditioned augmentation enriches local business contexts with artifacts of comparable activities. Across seven evaluated systems, reranking and iterative search improve answer quality, although full-corpus performance remains substantially below the gold-evidence reference. System rankings reverse between multi-hop and synthesis questions, revealing different strengths in relational search and information coverage. Scale experiments show lower performance as the company background expands, with questions and annotated supporting files held fixed. Additional analyses examine evidence use, computational costs, reader–judge sensitivity, and question and annotation quality through human assessment.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.