acceptodds
Under review as a conference paper at ICLR 2027

HeFLO: Flow-Regularized Offline Policy Learning under System Heterogeneity and Its Convergence Analysis

Abstract

Offline reinforcement learning (RL) learns control policies from pre-collected data without environment interaction, which is essential where online exploration is costly or unsafe. In practice such data are gathered from heterogeneous systems that share a dynamical structure but may differ in physical parameters. Additionally, the experts may themselves be specialized to particular configurations. We propose HeFLO: Heterogeneous FLow-regularized Offline policy learning, which minimizes an infinite-horizon average cost over a bounded range of parametric uncertainty with flow-based behaviour-cloning regularization. The objective is a domain randomization objective, and HeFLO is its offline, model-free, demonstration-regularized counterpart. For heterogeneous linear quadratic regulators we prove that a regularizer above an explicit and initialization-dependent threshold confines every iterate to the stability radius of the demonstrated behaviour, so the closed loop is stable for every configuration throughout training. Additionally, the iterates converge geometrically to a neighbourhood whose radius depends on the demonstration set. Empirically, HeFLO reaches near-expert performance on average over the configuration family across heterogeneous MuJoCo environments.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.