Decoding Trajectory Forgetting: Continual Reinforcement Learning for Diffusion Large Language Models via Trajectory Alignment
Abstract
Diffusion large language models (dLLMs) introduce a parallel decoding paradigm that resolves multiple masked positions with a flexible order. Recent developments in post-training for dLLMs have substantially improved dLLM reasoning capabilities in single-domain training and inference scenarios with reinforcement learning (RL). Yet, continually adapting dLLMs to evolving real-world tasks remains challenging, as previously learnt capabilities can rapidly degrade after model post-training on new domains. To investigate the reason behind catastrophic forgetting in dLLMs, we examine how previously learnt decoding trajectories evolve during continual RL from two complementary dimensions: (1) token prediction distributions that characterise what to decode, where continual RL can alter prediction behaviour on previous domains, and (2) decoding order patterns that characterise which positions to decode, where continual RL can distort the relative decoding order between token positions. Beyond changes in token predictions, we find that forgotten tasks exhibit more pronounced order drift in the decoding trajectory than preserved ones, revealing decoding trajectory forgetting as an important dimension of catastrophic forgetting in dLLM continual RL. Motivated by these insights, we introduce CLD, a novel continual RL framework for dLLMs that enables continual adaptation while preserving previously learnt decoding trajectories. During continual RL, the native RL objective drives adaptation to the incoming domain, while decoding trajectory alignment provides stability by preserving previously learnt decoding behaviour. Moreover, CLD can be easily integrated with different dLLM RL algorithms in a plug-and-play style. Extensive experiments on 9 benchmarks demonstrate that CLD effectively mitigates catastrophic forgetting in dLLMs. Our code is provided at https://anonymous.4open.science/r/CLD-2026
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.