Discrete Flow Matching for Offline-to-Online Reinforcement Learning
Abstract
Many reinforcement learning (RL) tasks have discrete action spaces, but most generative policy methods based on diffusion and flow matching are designed for continuous control. Meanwhile, generative policies usually rely heavily on offline datasets and offline-to online RL is itself challenging, as the policy must improve from new interaction without losing useful behavior learned from static data. To address those challenges, we introduce DRIFT, an online fine-tuning method that updates an offline pretrained continuous time Markov chain (CTMC) policy with an advantage-weighted discrete flow matching loss. To preserve useful pretrained knowledge, we add a path-space penalty that regularizes the full CTMC trajectory distribution, rather than only the final action distribution. For large discrete action spaces, we introduce a candidate-set approximation that updates the actor over a small subset of actions sampled from reference-policy rollouts and uniform exploration. Our theoretical analysis shows that focusing on a small set of high probability actions can reduce computation while keeping the approximation error small, making our method practical for large action spaces. We evaluate on macro-action MinAtar games where offline-pretrained DRIFT outperforms its online-only counterpart across all five games and achieves the highest mean return on three and on oracle free Jericho, DRIFT achieves a mean normalized score of 7.3% with a lightweight GRU encoder, outperforming the evaluated pretrained language-model baselines. Ablations demonstrate the importance of appropriately weighted path-space regularization and reference-guided candidate construction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.