Protecting Continual Reasoning with Depth-Informed Trust Regions
Abstract
Continual reinforcement learning with verifiable rewards requires reasoning language models to acquire new capabilities without erasing behaviors learned from earlier tasks. This setting is especially challenging when replay is disallowed: after a task ends, the learner must preserve useful prior behavior using only compact statistics rather than old prompts or trajectories. We introduce **DToP** (Depth-informed Trust Region over Parameters), a replay-free continual adaptation method that summarizes past reasoning behavior through complementary functional sensitivities and constrains future optimization within the resulting anisotropic parameter-space geometry. At each task boundary, DToP first identifies *late-settling* token states, whose intermediate-layer next-token distributions approach the final-layer prediction only at relatively deep layers. For these selected states, DToP measures the parameter sensitivity of the local decision margin between the sampled token and its strongest alternative, and compresses the resulting gradients into a low-rank second-moment memory. On the remaining candidate states, DToP estimates diagonal policy Fisher information to capture broader distributional sensitivity. The two memories are scale-normalized and combined into a joint trust-region metric. During subsequent tasks, DToP does not constrain optimizer steps independently. Instead, each proposed update is projected so that the *cumulative* parameter displacement from the previous task endpoint remains inside the boundary-centered trust region. This explicitly controls the accumulation of many individually small updates along historically sensitive directions. We evaluate DToP on a heterogeneous continual reasoning benchmark spanning multiple reasoning domains. The results provide encouraging evidence that DToP can better balance the acquisition of new capabilities with the retention of previously learned ones, while retaining only compact parameter statistics and requiring no inference-time modification.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.