Squared Family MDPS: Provably Efficient RL with Tractable Representations
Abstract
Regret bounds in continuous reinforcement learning require structural assumptions on the underlying MDPs. Many model-based parametric approaches are either difficult to theoretically analyse or hard to implement. We address this issue by proposing the Squared Family MDPs of which the transition kernels are parameterised by a family of tractable probabilistic models called squared families. Theoretically, we derive a regret bound, sub-linear in total time , linear in feature dimension and horizon length . Practically, we propose Representation RL with the Squared Neural Family (RESNEF), which uses Maximum Likelihood Estimation to learn the transition density and extract useful representations for downstream RL tasks. RESNEF brings up to 40% performance improvements compared with state-of-the-art baselines across diverse continuous control tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.