acceptodds
Under review as a conference paper at ICLR 2027

Agent Capability Shapes the Value of Post-Training Data

Abstract

Effective post-training of language-model agents depends on which experiences they learn from, as well as how much data they receive. We investigate whether an agent's current knowledge and capabilities can guide data selection for supervised fine-tuning (SFT) and reinforcement learning (RL). For SFT, we distinguish KNOWS and GAP traces by how accurately the agent predicts execution outcomes from known inputs. Selecting traces with larger environment prediction gaps yields a higher observed mean reward than random selection under matched training budgets. Cross-model comparisons show that this selection advantage depends on the learner. For RL, reward prediction gaps capture errors in the agent's estimates of its own rollout rewards. After controlling reward standard deviation, selecting high reward-GAP tasks improves final mean reward by relative to random selection, showing value beyond reward variability. We also identify practical proxies based on predicted rewards: traces and tasks selected by the proxies yield strong performances without additional verification. These findings connect the value of post-training data to the agent's predictive capabilities and offer a practical way to guide data selection. Code is anonymously available at https://anonymous.4open.science/r/agent_data-6BB4/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.