EXConan: A Global Benchmark and Evidence-Grounded Multi-Agent Framework for Extreme Weather Reasoning
Abstract
While recent large language models have demonstrated remarkable general reasoning capabilities, their potential in Earth system sciences remains bottlenecked by the complexity of meteorological tasks. Existing evaluations primarily target numerical forecasting, leaving a critical gap in assessing LLMs for evidence-grounded extreme weather reasoning that requires multi-source evidence interpretation and domain-specific knowledge. To address this, we introduce EXConan-Bench, a comprehensive global benchmark comprising 8,498 expert-curated instances. By integrating multi-source event records with ERA5 meteorological reanalysis data, the benchmark covers eight extreme weather hazards and normal conditions, challenging models to identify events using structured atmospheric evidence and interpretable rationales. Furthermore, we propose \conans EXConan, an evidence-grounded multi-agent reasoning framework designed to overcome the limitations of direct prompting. EXConan introduces an agentic paradigm that meticulously coordinates specialized meteorological experts, evidence-based candidate evaluation, and counterfactual verification to ensure reliable decision-making. Extensive experiments across representative closed-source and open-source LLMs reveal the significant limitations of current baselines in meteorological reasoning. Remarkably, our EXConan framework effectively mitigates these shortcomings; it improves Qwen3-30B from 49.33% to 54.63%, with consistent gains in Macro Precision, Weighted-F1, and Kappa. Ultimately, EXConan establishes a scientifically grounded paradigm for reliable weather intelligence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.