Zephyrus-R1: Agentic Reinforcement Learning for Weather Reporting
Abstract
Reinforcement learning (RL) has emerged as a key method for training AI models in scientific applications, but RL remains difficult when outputs cannot easily be verified against a known ground-truth answer. Natural-language weather forecast reports provide one such valuable domain example: these reports translate atmospheric data into assessments of weather evolution, hazards, and uncertainty, and they are valuable both to the general public and key decision makers. Unlike numerical weather-analysis questions, they admit many valid outputs, and deterministic verification is inherently limited due to the open-ended, natural-language setting. To solve these challenges, we develop Zephyrus-R1, a practical framework for RL on scientific weather forecast reporting. We focus on three key components of the agentic RL pipeline: reward design, curriculum learning, and an analysis of human-expert alignment. To encourage both report legibility and factual accuracy, we design a reward function that combines a meteorological rubric with bidirectional claim entailment, both evaluated against National Weather Service reference reports. Our framework optimizes this reward in a simulated tool-calling environment, enabling vision-language agents to gather and analyze forecast evidence across regions and lead times before producing a report. We use annotations from three meteorologists to audit the claim labels produced by our LLM judges. Finally, we introduce a curriculum learning process to help small models first develop tool use and scientific reasoning through RL on verifiable weather-analysis questions, then optimize the open-ended report reward directly. After training Qwen3.5-4B with Zephyrus-R1, we improve the mean across six judge–metric scores by 26% over the base model and achieve results competitive with frontier open-source models. Meanwhile, after training Qwen3.8-27B, we improve the same mean by 13% over the base model, by 6.7% over Kimi-K3, and by 2.7% over GPT-5.6-Luna.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.