Theoretical Insights into Informed Value Estimation for Asymmetric Recurrent Natural Actor-Critic
Abstract
In partially observable environments, reinforcement learning agents must act based on incomplete or noisy observations. While asymmetric actor-critic methods empirically improve learning, their theoretical properties remain largely unexplored. We generalize the Recurrent Natural Actor-Critic (Rec-NAC) algorithm to the asymmetric setting, where a state-dependent privileged signal is available only during training and is used by the critic. Both actor and critic consist of a recurrent encoder of the observable history followed by a linear head; the critic's head is learnable and integrates the privileged signal, while the actor does not. By adapting the finite-time and finite-width analysis of policy evaluation via recurrent temporal-difference learning to this setting, we identify distinct information and representation benefits of an informed critic and uncover a fundamental trade-off: privileged information reduces conditional target variance and removes the irreducible error due to information aliasing, while the informed architecture can improve best-in-class approximation but also induces additional finite-time complexity. We further analyze the resulting recurrent natural policy-gradient update and show that privileged signals induce an information mismatch in the actor's compatible function regression. Our results characterize when privileged information improves value estimation and its implications for policy learning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.