Advanced Policies: A First-Principles Path from Policy Gradient to Q-Learning
Abstract
We apply entropy to reinforcement learning along two axes: entropy-augmented reward defines the soft objective, while relative-entropy regularization controls the size of a policy-improvement step without changing that objective. Together they yield a one-parameter family of advanced policies connecting a base policy to its soft-greedy policy. Under exact evaluation, every nontrivial member improves upon the common base, although the improvement need not be monotone along the family. With a consistent policy and action-value function, this family forms an entropic mirror-descent path. We then relax this consistency, treating the policy and action-value function as independent coordinates of the advanced policy. Differentiating the objective of the advanced policy yields Advanced Actor–Critic (AAC), whose endpoint gradients recover soft policy gradient and an action-centered Q-learning-like update. We develop implementations for discrete and continuous actions. Experiments validate both endpoints and demonstrate effective learning at intermediate parameter values.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.