FluForecastBench: Influenza Forecasting with Publication-Time Information
Abstract
Reliable influenza forecasting requires accurate predictions and useful uncertainty estimates from surveillance reports that are delayed, incomplete and revised. Retrospective evaluation using later reports can conceal the information constraints faced by a forecaster. Although prior studies establish the importance of reporting revisions, comparisons must also account for missing observations and changes in the data used to calibrate predictive uncertainty. To make these constraints explicit in evaluation, we introduce FluForecastBench, covering four countries, eight tasks and 90 surveillance series. It pairs histories selected under explicit publication rules with specified later editions and evaluates both forecasts against the same targets. The US condition uses archived early reports. Other tasks follow their recorded publication or first-seen rules. We evaluate twelve baseline procedures across seven per-task leaderboards. Relative performance varies across tasks: seven fitted procedures have higher error than a seasonal naive forecast on the German state panel. A separate five-baseline experiment in the United States and Japan separates model selection from interval calibration over two seasons. All four fitted baselines improve mean point accuracy under publication-constrained inputs, but nominal 80% intervals cover only 39–54% of US targets and 15–40% of Japanese targets under a shared preceding-season residual calibration scheme. Further analyses distinguish revised values from differences in available history and calibration data. These results identify incomplete histories and the transfer of uncertainty calibration across seasons as concrete challenges for influenza forecasting.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.