acceptodds
Under review as a conference paper at ICLR 2027

AC/DT: Actor-Critic Decision Transformer For Offline-to-Online Reinforcement Learning

Abstract

We introduce the Actor‑Critic Decision Transformer (AC/DT), an offline‑to‑online (O2O) reinforcement learning transformer-based method that replaces the behavior‑cloning objective of Decision‑Transformer‑style models with a value‑based actor‑critic optimization over a sequence‑conditioned stochastic Transformer policy. AC/DT combines long-horizon sequence conditioning with explicit policy improvement through twin Markovian critics. To exploit the temporal information available within each sampled context, the critics are trained position-wise using context-truncated -returns, providing multi-step value targets while retaining lightweight Markovian value functions. During offline pretraining, we mitigate out-of-distribution overestimation through conservative value regularization. During online fine-tuning, conservative regularization is restricted to offline experience. On D4RL benchmarks, AC/DT moves beyond the offline data distribution and achieves competitive performance against action‑likelihood‑based Transformer baselines such as ODT, hybrid variants like ODT+TD3 and other state-of-the-art O2O methods. A state-matched distributional analysis further shows that AC/DT reduces adherence to the dataset action marginals relative to likelihood-based ODT, providing evidence of a consistent shift in the policy-induced action marginals toward an expert reference.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.