Historical Information in Forecaster Selection: Replay, Retention, and Candidate Exclusion
Abstract
Historical archives contain reports unavailable when forecasts would have been issued. These reports can change both candidate rankings and which checkpoints survive for final selection. We vary the information available at retention and final selection while holding fitted trajectories fixed and testing selected forecasters on common deployment inputs. Our main study predicts vehicle positions 60 seconds ahead across eight refitted episodes at two transit agencies, SEPTA and VTA. Retrospective selection increases position error in seven episodes. At equal additional realistic evaluation counts, retaining runner-up checkpoints and replaying fewer dates improves mean error over complete replay of one winner per group in six episodes, with agency-average gains of 4.48 m at SEPTA and 0.10 m at VTA. Restricting the library to origin-available fits remains competitive. Finitecandidate analysis separates exclusion, replay-score error, and changes in candidate performance between validation and testing. Full-replay controls distinguish the benefit of wider retention from the cost of reduced date coverage. Accurate rescoring cannot recover a discarded alternative, while retaining more candidates need not improve deployment loss. Retention and final selection must therefore be evaluated together.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.