Edit as Action: Reinforcement Post-Training for Time-Series Forecasting
Abstract
Even a trained forecaster can mistime peaks or damp oscillations in an otherwise plausible trajectory. Correcting these features is not a local shape-matching task: an edit succeeds only if its complete forecast improves. We introduce EditCast, an outcome-guided semantic RL post-training framework built on the edit-as-action principle. It is a one-step, full-information contextual RL problem: training futures score each candidate through full-horizon mean squared error (MSE) and structural rewards, while a lightweight policy selects a semantic action from the history and reference forecast. The finite semantic action bank makes expected reward exactly computable. Across 308 checkpoint-level evaluations spanning eight datasets, four architectures, and four horizons, validation-selected EditCast reduces aggregate MSE and mean absolute error (MAE) by 1.1% and 0.8% relative to the frozen checkpoints, without updating backbone parameters. A fixed structural-reward weight raises joint accuracy–structure wins from 31 with accuracy-only editing to 59; validation selection reaches 64. These results establish outcome-guided semantic RL as a post-training stage for frozen forecasters.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.