acceptodds
Under review as a conference paper at ICLR 2027

SPECTRA: A Multi-Scenario Benchmark for Granular RAG Evaluation in Chemistry

Abstract

Despite rapid progress in large language models (LLMs) and agentic systems, retrieval-augmented generation (RAG) remains indispensable in any setting where authoritative answers must come from a controlled, often proprietary corpus such as scientific literature, internal company documentation, or regulatory archives. We introduce (cenario-artitioned xpert hemistry estbed for AG ssessment), a benchmark for fine-grained evaluation of RAG systems on expert-level questions grounded in open-access scientific publications. The benchmark organizes 2007 questions across 8 chemistry domains into a two-dimensional taxonomy: 6 question scenarios (single-hop extraction, multi-hop reasoning, conditional filtering, comparison, aggregation, and multi-turn dialogue) that isolate distinct system capabilities and 3 orthogonal modification types (multimodal, constrained, and negative queries). We also propose an evaluation framework to assess LLMs and RAG configurations across several metrics and apply it to 7 RAG configurations and 3 standalone LLM baselines, with the primary Answer Accuracy judge validated against two experts on 200 usable examples. The benchmark evaluation yields two key findings. First, scenario- and modification-level evaluation on one domain reveals performance trade-offs invisible to aggregate scoring: configurations that appear equivalent by overall accuracy differ substantially across question types. This allows practitioners to optimize RAG configurations according to their primary use case. Second, the dedicated multi-turn scenario illustrates how per-turn analysis supports fault localization: document-level DOI Recall remains stable across dialogue turns for hybrid configurations, while evidence-level Source Support declines modestly and Answer Accuracy declines more strongly. This shows that degradation is not explained by complete failure to retrieve the source article. While the corpus is chemistry-focused, the scenario taxonomy and evaluation protocol are not intrinsically chemistry-specific and can be adapted to other scientific corpora. Beyond evaluation, the per-scenario performance vectors produced by our framework can support data-driven RAG configuration design, and the semi-automated question generation pipeline can suggest a route toward periodic question refresh to reduce contamination. The dataset and evaluation code are publicly available through anonymized repositories: https://huggingface.co/collections/anonymousauthor2026team/scientific-rag-benchmark-collection and https://anonymous.4open.science/r/1c7ffbd3c15e.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.