acceptodds
Under review as a conference paper at ICLR 2027

Auditing Tabular Foundation Models for Wildfire Forecasting: Transfer, Local Adaptation, and Weather-Response Stability

Abstract

Wildfire forecasts must rank rare fire days from weather predictions, often before a region has local fire labels. We evaluate TabPFN and TabFM against CatBoost, LightGBM, logistic regression and the Fine Fuel Moisture Code (FFMC), using complete fire journals from two Russian regions and held-out seasons in British Columbia (BC). Neither foundation model beats its selected classical comparator in any of six zero-label transfer contrasts; FFMC recalls 3–5 percentage points more fires at a 5% alert budget. With 64 local BC labels, averaging five contexts improves both foundation models, but averaged trees reach the same recall. On BC 2023 the foundation models exceed the classical family selected by single-context validation by about 0.9 average-precision points and 1.8–2.0 recall points. The prespecified rule for confirming this advantage on untouched BC 2024 fails: only 2 of 16 simultaneous intervals are positive, although context averaging still helps. Weather-response categories disagree across contexts for 27–28% of queries at 4,096 labels, compared with 17% for logistic regression. Agreement across contexts selects FFMC-teacher response directions no better than response magnitude at matched coverage. Ranking, response stability and FFMC-equation fidelity therefore require separate evaluation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.