EnterpriseRAG-Bench: A Realistic Benchmark for Cross-Document and Multimodal RAG over Medium-Scale Corpora
Abstract
Retrieval-augmented generation (RAG) evaluation has focused heavily on short-context question answering, synthetic retrieval tasks, and single-document reasoning. These settings underrepresent realistic enterprise workloads, where systems must retrieve, preserve, and synthesize evidence across long documents, noisy scans, audio, presentations, and scientific visual media. We introduce EnterpriseRAG-Bench, a medium-scale benchmark for end-to-end evaluation over one heterogeneous corpus. The current release contains 325 scored questions in five 65-question suites: factual, comparison, analytical, multi-hop, and multimodal. Of these items, 303 are answerable and 22 are unsupported-by-design abstention probes. The multimodal suite contributes 45 evidence-supported questions over presentation video and slide content, scanned handwriting, audio, and scientific video/frame analysis, plus 20 abstention probes. The benchmark combines human-verified gold answers with semantic judging, deterministic factual-anchor checks, source annotations, and media-specific audit metadata. Five systems have results across all five suites, while five externally executed Open WebUI baselines currently cover the four text suites. The results expose substantial variation by task and access regime: VaultIQ has the highest five-suite mean at 67, Claude Opus 5 has the highest simple-factual score at 88, VaultIQ and ChatGPT 5.6 Sol tie on multi-hop at 62, and VaultIQ has the highest multimodal score at 72. A fixed-pipeline Open WebUI study also reproduces prior findings that generator scale alone is not a reliable proxy for end-to-end RAG performance. We release the benchmark specification, scoring rubric, source manifest, and evaluation harness to support reproducible extension.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.