acceptodds
Under review as a conference paper at ICLR 2027

Squared Family MDPS: Provably Efficient RL with Tractable Representations

Abstract

Regret bounds in continuous reinforcement learning require structural assumptions on the underlying MDPs. Many model-based parametric approaches are either difficult to theoretically analyse or hard to implement. We address this issue by proposing the Squared Family MDPs of which the transition kernels are parameterised by a family of tractable probabilistic models called squared families. Theoretically, we derive a regret bound, sub-linear in total time , linear in feature dimension and horizon length . Practically, we propose Representation RL with the Squared Neural Family (RESNEF), which uses Maximum Likelihood Estimation to learn the transition density and extract useful representations for downstream RL tasks. RESNEF brings up to 40% performance improvements compared with state-of-the-art baselines across diverse continuous control tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.