Train with Hindsight: Post-hoc Checkpoint Selection atop Deterministic Computing
Abstract
The training of large language model(LLM) involves substantial steps, yet the optimal checkpoint for evaluation or next training stage can exist inside the training trajectory rather than at its endpoint. Fixed-interval checkpointing couples the set of selectable states to a priori what to save before that trajectory is observed, while denser storage and evaluation increase checkpoint I/O, storage, and validation costs. We introduce Post-hoc Checkpoint Selection (PhCS), a ”train with hindsight" framework that decouples checkpoint selection from physical checkpoint storage. In the original training, PhCS saves parameters and optimizer states periodically as anchors and records lightweight trajectory metadata and input per step, after which post-hoc selection policies identify candidate states on the realized trajectory. Atop deterministic computation, reproduction of candidate states begins from the nearest preceding anchor and replay the previous input along the same numerical path. We instantiate PhCS in SFT-RL pipeline: static selection in SFT covers the entire trajectory, combining loss, token accuracy, and entropy, whereas dynamic selection in GRPO conducts recursive exploration on validation set. PhCS thereby turns checkpointing from a precommitted storage schedule into a post-hoc decision over the realized training trajectory. We evaluate PhCS in ReTool framework, a representative SFT-Agentic RL scenario; the experiments demonstrate that PhCS achieves higher pass@4 than conventional training paradigms on multiple datasets such as AIME24/25/26 and AMC23 , and achieves superior equivalent storage efficiency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.