EarthDataBench: Benchmarking LLM Agents for Long-Horizon Scientific Data Production
Abstract
Transforming heterogeneous Earth data into reliable scientific products requires complex, labor-intensive processing, limiting the pace and scale of research. Although large language model (LLM) agents have demonstrated success in automating a range of scientific tasks, their ability to produce research-quality Earth-science data remains insufficiently assessed. We introduce EarthDataBench, a benchmark of 40 executable, paper-derived tasks spanning seven Earth-science topics and diverse data sources, processing mechanisms, and product types. Tasks require agents to complete long-horizon scientific workflows, integrating heterogeneous data and implementing domain-specific methods to produce research-quality datasets. Unlike existing QA benchmarks that rely on matching reference answers, our evaluation combines comparisons with expert products and validation against independent observations, enabling assessment of whether agent-generated products match or surpass expert-reference quality. Across nine LLMs and ten model–harness configurations, 52.0% of outcomes pass scientific-consistency checks, but only 19.3% reach expert-reference quality or higher and 0.8% outperform expert references on independent validation. Our empirical results suggest that reliable execution remains a key bottleneck in scientific data production, even when methods are prescribed. EarthDataBench provides a testbed for advancing research-ready data production in Earth science and assessing improvements beyond expert references.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.