Not All Experiences Are Equal: Counterfactual Experience Valuation for Multi-Turn Agent RL
Abstract
On-policy reinforcement learning for multi-turn language agents largely optimizes the update rule (advantages, group comparisons, KL) while treating which experiences deserve update mass as a default. Return and advantage record how an episode ended and how strongly the current surrogate wants to reinforce it. Neither answers the counterfactual that matters under a conserved budget: whether giving one experience more update mass, taken from its neighbors, improves held-out capability. We introduce EVA (Experience Valuation and Augmentation), a data-centric layer in front of an unchanged , and its measurement core IA-ELV. IA-ELV scores each experience by a budget-preserving probe that returns gold coordinates , namely meta-performance change and decision-uncertainty reduction. On ALFWorld and WebShop, at 1.5B and 7B, learning value is sparse and peaks at mid-training, where return and advantage fail to recover . A leak-free predictor ranks the gold object with the same sign, so EVA can sit in front of GiGPO without a probe at every step. Gold Mid allocation, and the same rule under Mid-continue GiGPO, beat uniform weighting across both environments and both scales; Mid is the strongest stage. Soft probes keep learning value distinct from augmentation gain. Legally recollected four-action routing then adds held-out value beyond reweighting alone. EVA therefore separates outcome, credit, learning value, and augmentation value under explicit conservation constraints. Our code is available at https://anonymous.4open.science/r/eva-3B8C.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.