Forecasting as a Verifiable Task: What RL and Test-Time Compute Can (and Cannot) Buy
Abstract
Probabilistic forecasting is verifiable: once the future arrives, a proper scoring rule grades every sampled trajectory. Time-series foundation models are therefore natural targets for the reinforcement-learning (RL) post-training and test-time-compute recipes that transformed language models. We ask what these recipes buy when the verifier grades a distribution rather than a single answer. We decompose the excess CRPS of a sampled forecaster into Monte-Carlo, optimisation and information gaps, derive the per-sample reward under which policy gradient is unbiased for CRPS (it is necessarily pairwise), and analyse GRPO under this reward. GRPO's group-relative advantages are provably biased for pairwise rewards: the group-mean baseline lowers the effective dispersion coefficient from 1/2 to (G-2)/(2(G-1)), and per-group standard-deviation scaling adds an O(1) shrinkage that no group size removes. The bias predicts the under-dispersion of every uncorrected run across two model families and three model sizes (coverage 0.63 to 0.55 for Chronos-T5-small); a two-line correction removes it. Corrected RL matches a supervised fine-tuning control at equal budget, with no significant GIFT-Eval gain, while preserving calibration at half the KL divergence. Test-time compute fails for a different reason: its only verifier, a self-backtest, has fidelity rho ≈ 0.1, so every backtest-driven intervention hurts. The bias and its correction apply to any pairwise or set-level reward. Code is available at https://anonymous.4open.science/r/verifiable-forecasting-72ED.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.