Order-RFT: Gradient Alignment and Age-Aware Data Selection for Efficient Reinforcement Fine-Tuning
Abstract
Reinforcement fine-tuning (RFT) improves the reasoning ability of large models, but rollout generation dominates its training cost. Data selection lowers this cost by prioritizing high-utility score samples, yet it faces an efficiency–adaptivity trade-off. Cached-score methods—whether static or dynamically updated—are cheap, but they lag behind the evolving utility in RFT. This risks a stale-utility exclusion problem: samples with low cached scores may remain unselected even after becoming useful, leaving their scores uncorrected. Conversely, refreshing scores for all samples consumes the rollouts that selection aims to save. To tackle this problem, we propose Order-RFT, a plug-and-play framework that estimates utility before rollout generation while tracking evolving utility through scheduled re-examination. It combines two complementary components: (i) a historical gradient-alignment estimator that approximates each sample’s marginal contribution to policy improvement; and (ii) a linear age bonus mechanism that prioritizes both cached utility and time since the last observation, enabling samples with cached scores to be reassessed. We establish approximation error bounds for the estimator and re-examination guarantees for the age mechanism. Experiments on several benchmarks show that Order-RFT achieves higher matching with refreshed sample-utility rankings and outperforms both static and dynamic baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.