Does Your Wildfire Prediction Model Actually Work, or Just Score Well?
Abstract
Earth foundation models can be transferred to wildfire prediction, but their reported downstream scores depend on the evaluation choices used to compute them. We formulate wildfire transfer as a fixed-contract evaluation problem and test how evaluation choices affect scores and model rankings. Holding predictions fixed, decision changes from to solely with the matching rule; holding frozen features fixed, changing the head-selection metric changes decision by up to percentage points. We introduce the contract dominance ratio and rank instability rate, showing that evaluation choices affect scores more than backbone choice and reverse the relative ranking of up to 46% of model pairs across configurations. We develop WILDFIRE-FM to provide reference performance across six downstream tasks under specified contracts. Under matched contracts, this wildfire reference leads the evaluated frozen backbones on occupancy, spread, and burned-area prediction. The reference provides a baseline for comparing Earth-FM transfer across tasks, while the controlled checks show why these comparisons require explicit evaluation contracts. See implementation details at: https://anonymous.4open.science/r/Wildfire-fm-evaluation-contracts-5AE9.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.