EarthArena: Benchmarking Agents on Real-world, Long-Horizon Scientific Tasks in Earth Observation
Abstract
Scientific investigations in Earth observation combine heterogeneous observations, geospatial processing, and interdisciplinary analyses into long-horizon workflows with interdependent stages. Agents must apply scientific methods appropriately and maintain the scientific validity of intermediate artifacts and final conclusions. However, whether they can meet these requirements throughout the workflow remains to be systematically evaluated. To this end, we introduce EarthArena, an executable benchmark for real-world, long-horizon scientific tasks in Earth observation. We construct 101 tasks from 30 peer-reviewed publications spanning diverse areas of Earth observation. Each case organizes tasks linked by intermediate artifacts around a scientific objective, with reproduced reference results and human-reviewed scoring rubrics. EarthArena links scientific conclusions to supporting artifacts and execution evidence to assess agents' ability to maintain scientific validity throughout the workflow. Across six language models and three agent harnesses, the best-performing configuration achieves a case-averaged Rubric score of 44.3. Model rankings vary across harnesses, highlighting the importance of evaluating scientific task performance at the level of complete model-harness configurations. EarthArena provides a testbed for systematically evaluating and improving agents' ability to conduct real-world, long-horizon scientific investigations in Earth observation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.