acceptodds
Under review as a conference paper at ICLR 2027

Order-RFT: Gradient Alignment and Age-Aware Data Selection for Efficient Reinforcement Fine-Tuning

Abstract

Reinforcement fine-tuning (RFT) improves the reasoning ability of large models, but rollout generation dominates its training cost. Data selection lowers this cost by prioritizing high-utility score samples, yet it faces an efficiency–adaptivity trade-off. Cached-score methods—whether static or dynamically updated—are cheap, but they lag behind the evolving utility in RFT. This risks a stale-utility exclusion problem: samples with low cached scores may remain unselected even after becoming useful, leaving their scores uncorrected. Conversely, refreshing scores for all samples consumes the rollouts that selection aims to save. To tackle this problem, we propose Order-RFT, a plug-and-play framework that estimates utility before rollout generation while tracking evolving utility through scheduled re-examination. It combines two complementary components: (i) a historical gradient-alignment estimator that approximates each sample’s marginal contribution to policy improvement; and (ii) a linear age bonus mechanism that prioritizes both cached utility and time since the last observation, enabling samples with cached scores to be reassessed. We establish approximation error bounds for the estimator and re-examination guarantees for the age mechanism. Experiments on several benchmarks show that Order-RFT achieves higher matching with refreshed sample-utility rankings and outperforms both static and dynamic baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.