EarthLitBench: Benchmarking Literature-Grounded Scientific Inquiry in Earth Science
Abstract
Scientific inquiry in Earth system science demands synthesizing fragmented, cross-study empirical findings across diverse environmental spheres and temporal-spatial scales. Evaluating LLM support for this work therefore calls for assessing evidence selection and scientific justification together. However, existing benchmarks predominantly focus on single-document factoid QA, leaving a critical blind spot in evaluating multi-hop reasoning and evidence synthesis across extensive literature. To bridge this gap, We introduce EarthLitBench, a benchmark of 822 instances across ten tasks organized into four capability groups: problem formulation and literature discovery, study and evidence understanding, scientific reasoning, and knowledge synthesis. Grounding our benchmark in authoritative scientific reviews, we construct traceable knowledge graphs that capture complex scientific relationships, thereby guiding rigorous question generation and evidence verification. To ensure fair evaluation, source reviews, construction graphs, and reference solutions are strictly withheld from candidate models, while evidence retrieval is constrained within an authorized reference corpus. Furthermore, we establish the RERA evaluation protocol, which couples automated source verification with reference-guided LLM judges across four complementary dimensions: Retrieval, Evidence fidelity, Reasoning fidelity, and Answer fidelity. Experiments evaluate nine LLM-only models and compare four of them with fixed retrieval. Models perform relatively well on problem decomposition but have greater difficulty with scientific reasoning, even though explicit references are given. These findings highlight a gap between accessing scientific literature and conducting scientific inquiry, motivating stronger integration of evidence and scientific interpretation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.