T-MAC: Target-Motion Actor-Critic for Generative Reinforcement Learning
Abstract
Generative policies can represent complex and multimodal action distributions, which makes them attractive for continuous control. However, in online reinforcement learning, the target of such a policy is defined by a critic that changes after every update, so the policy keeps chasing a moving target. Existing methods train the policy toward where the target is now, but do not use how the target has moved. We propose Target-Motion Actor-Critic (T-MAC), a flow-based actor-critic method that learns from both signals. In addition to the current-target loss, T-MAC trains the cross-update change of the policy velocity field to follow the target motion. T-MAC estimates the target motion by comparing the Jordan–Kinderlehrer–Otto (JKO) responses of successive critics on the same source actions, and uses a response-magnitude gate to weight each state by how far its target moves. Experimental results on six MuJoCo and four HumanoidBench tasks show that, compared with representative model-free and generative-policy RL baselines, T-MAC achieves state-of-the-art performance with consistent improvements across all ten tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.