From AI Forecast to Text Report: Pangu-Driven Weather Report Generation via Vision-Language Fine-tuning
Abstract
Recent work, WeatherSyn showed MLLMs can generate them from ERA5 reanalysis images, but ERA5 has multi-day latency. More fundamentally, it is produced via 4D-Var assimilation over a 12-hour window thus using it as model input for real-time forecasting can leak future information and inflate apparent skill. Another work, WeatherQA, by contrast, targets severe-weather reasoning and text-only discussion generation, not forecast heatmaps, thereby omitting critical spatial information, limiting operational utility. In this paper, we ask whether AI-based weather forecast fields, specifically from Pangu-Weather, can replace ERA5 as the visual input for driving text report generation. We construct an end-to-end pipeline that renders city-level meteorological heatmaps for 31 U.S. cities, pairs them with professional NWS forecast discussions from the WSInstruct corpus, and LoRA fine-tunes Qwen2-VL-7B on 2,016 paired samples from 2020–2021. Our central contribution is a controlled input-source ablation: we train two models under identical hyperparameters, differing only in whether the visual input is rendered from Pangu forecast fields or ERA5 reanalysis fields, and evaluate both on both test sets. On 1,043 test samples from 2022, the input-source cost of replacing ERA5 with Pangu is modest—at most ROUGE-L and BLEU-1—and training on Pangu forecast fields does not hurt downstream report generation. These results indicate that end-to-end, real-time AI weather forecasting pipelines, from initial conditions to human-readable reports, are feasible without relying on delayed reanalysis products. This is the first systematic study of whether AI forecast fields can replace reanalysis in driving text weather report generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.