acceptodds
Under review as a conference paper at ICLR 2027

Anisotropic Support-Transport Robust Offline Model-based Reinforcement Learning

Abstract

Offline model-based reinforcement learning (MBRL) suffers from a structural failure mode: the learned dynamics model is accurate on the data support but drifts rapidly off-support, where the policy naturally explores. Existing remedies—uncertainty penalties, pessimistic surrogate MDPs, and adversarial model learning—reduce trust to a single scalar quantity and commit to a fixed rollout horizon, ignoring the anisotropic geometry of real offline datasets. We introduce **ASTRO** (Anisotropic Support-Transport Robust Offline RL), which replaces the scalar trust radius with a data-adaptive *support tube* carved by local *k*-nearest-neighbor covariance and density estimates. The same geometric primitive gates both a robust value-aware model-learning loss—driving predicted latent values toward the minimum over the tube—and a state-dependent rollout horizon selected by a learned trust score. On 24 D4RL datasets spanning the Adroit and MuJoCo suites, ASTRO consistently matches or surpasses strong model-based baselines, with the largest gains on support-fragmented datasets where scalar trust abstractions are least appropriate.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.