Functional Critics for Off-Policy Actor-Critic: From Stability to Exploration
Abstract
Actor-critic methods are central to modern reinforcement learning, yet off-policy actor-critic theory remains far removed from practical algorithms, due in part to the interaction between the deadly triad and the continually changing target policy. We revisit functional critics, i.e., value functions that explicitly condition on the policy, and show that this representation provides a shared foundation for off-policy training stability and policy-space exploration. We first develop a convergent off-policy actor-critic framework under linear function approximation that accommodates target-based TD, partial coverage, and evolving behavior policies. Our analysis combines a policy-evaluation guarantee under partial coverage with policy-gradient generalization of functional critics to yield a dual trust-coverage framework that controls policy-gradient error under evolving behavior policies while enabling greater reuse of off-policy data. We further develop a replay-buffer analysis under i.i.d. sampling, yielding a principled replay-sampling design for minimizing the replay-sample complexity required to reach a prescribed evaluation accuracy. We next show that functional critics provide a natural representation for model-free posterior-sampling-style exploration. Standard critic ensembles do not naturally represent joint uncertainty over return functions across policy space, whereas functional-critic ensembles represent policy-conditioned return functions directly. This motivates a particle-based, model-free analogue of posterior-sampling-style exploration without explicit model learning. Finally, we instantiate these ideas in a deep functional actor-critic algorithm that, despite omitting several standard off-policy AC heuristics, achieves strong performance across the full Dog and Humanoid task series of the DeepMind Control Suite.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.