PADGen: Policy-Adaptive Data Generation for Offline Reinforcement Learning
Abstract
Data augmentation has been widely used to improve offline reinforcement learning (RL) by expanding the available training data. However, existing methods largely follow a static generation paradigm, where augmented data remain fixed throughout policy optimization, overlooking that data utility changes as the policy evolves. Consequently, statically generated data may become less useful as training progresses and fail to accommodate the evolving data needs of the policy, thereby limiting further policy improvement. To address this limitation,we propose a policy-adaptive data generation framework (PADGen), which consists of a generator, selector, and policy. The generator generates data based on the utility feedback from both the selector and the policy. The selector evaluates and selects generated data according to their utility for policy optimization. After each training iteration, the policy provides optimization feedback, while the selector returns the utility evaluation of the generated data. The generator then leverages these feedback signals to adaptively generate new data for the next iteration, forming a closed-loop data generation and policy optimization process. This design enables the generated data to be continuously adjusted according to the current policy, improving their utility throughout policy optimization. Extensive experiments across multiple offline RL algorithms and benchmark datasets demonstrate that PADGen consistently outperforms existing data augmentation methods. Our code is available at https://anonymous.4open.science/r/PADGen-ABCD.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.