Virtual Meteorologist’s Exam: Benchmarking Meteorological Expertise in Large Language Models
Abstract
Recently, the growing demand for virtual meteorologist assistants in operational scenarios has driven the incorporation of meteorological corpora into the training of large language models (LLMs). However, due to the scale and diversity, existing benchmarks fail to provide a comprehensive assessment of LLMs’ understanding and reasoning capabilities in meteorology. To this end, we introduce a benchmark, namely Virtual Meteorologist's Exam (VME), to evaluate whether existing large language models have professional-level meteorological expertise equivalent to that of humans. To be concrete, our VME comprises nearly 6,500 question–answer pairs across four formats, making it four times the size of comparable datasets. It covers both academic knowledge and operational scenarios. Meanwhile, it is also organized into 5 primary categories and 25 fine-grained subcategories, enabling a detailed analysis of existing models’ strengths and weaknesses across meteorological subdomains. Based on that, we employ numerous state-of-the-art large language models, either general-purpose or meteorologically specialized, to examine their capability ceilings and preferences. Our results reveal critical weaknesses in current LLMs across specific meteorological subdomains, particularly meteorological observation, indicating key directions for further improvement. We hope that VME will serve as a standardized testbed for tracking progress and guiding the development of meteorological LLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.