Extrema: Enabling and Evaluating Agentic Research on Extreme Weather
Abstract
Understanding extreme weather is vital for climate risk management, but it demands scarce domain expertise. This barrier deepens global inequalities in climate resilience. AI agents offer a path toward democratizing such research. However, existing agents struggle with long-horizon inquiry, as error-free code cannot guarantee physical validity. We introduce Extrema, a system combining ExtremaWorld—which unifies 20 datasets and 100 validated tools—with ExtremaAgent, which pairs a scientific review loop and a persistent task graph to guard against methodological errors. To evaluate such investigations, we introduce ExtremaBench, which distills 58 tasks from 19 published studies covering five extreme weather categories. Withholding the source study, we require each agent to rediscover the study's findings independently, and grade the resulting conclusions and supporting figures against expert-authored rubrics. Powered by GPT-5.6-sol, Extrema scores 71.58/100, outperforming the strongest baseline by 10.07 points.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.