acceptodds
Under review as a conference paper at ICLR 2027

DataLens: Predicting Finite-Step Data Utility for Data-Efficient LLM Reinforcement Learning

Abstract

Selecting effective training data is crucial for efficient reinforcement learning (RL) of large language models (LLMs), yet example utility cannot be reliably inferred from local training signals. Existing gradient-based methods use local gradient information as a proxy for training improvement. We show that this approximation can break down for finite displacements: candidates with similar local alignment may produce substantially different target-objective improvements. We formulate finite-step data utility prediction and define utility as the target-objective improvement induced by a prescribed candidate-specific update. We propose DataLens, which augments first-order alignment with directional objective geometry derived from a finite-step expansion, modeling objective evolution along candidate-induced update directions. Controlled matched-objective experiments verify that this geometric signal captures utility differences missed by local approximations. Across six model-domain settings spanning Qwen, DeepSeek, and Mistral on math and code, DataLens achieves the highest point-estimate correlation with finite-step utility and transfers to short-horizon GRPO. In end-to-end GRPO experiments, 100 examples selected by DataLens achieve 84.0% MATH accuracy across three random seeds, compared with 70.1% using first-order alignment and 67.5% using random selection. These results show that RL data selection benefits from modeling finite-step objective geometry beyond local gradient signals.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.