Policy Improvement and Preservation in Continual Reinforcement Learning with Latent Families
Abstract
Continual reinforcement learning requires agents to improve in recurring environments while preserving performance on environments that are currently unavailable. Sharing policies enables reuse, but updates for one environment can harm others. We introduce a continuous-state framework that separates identification of latent environment families from the decision to share a policy. Learned Bellman comparisons identify families, while one-sided certificates use finite stored data to bound loss relative to retained reference policies, called anchors. When a shared update cannot be certified, policy splitting allows learning to continue. Under stated separation, sampling, coverage, and local neural-geometry assumptions, we establish high-probability guarantees for correct identification, anchor-relative preservation, and final -optimality, using at most one deployed policy module per family. The complete observation bound retains coverage, conditioning, representation-complexity, and acquisition factors. Under the stronger specialization that these factors remain uniformly controlled, the prescribed exploration schedule gives sufficient observation complexity where is the number of families, is module capacity, counts segments, and is a diagnostic family separation lower bound. Under additional compatibility and critic-accuracy conditions, all families can remain in one deployed module without splitting; matched constructions show when separate policies are necessary. This module reduction does not imply lower observation or total-memory cost. In a recurrent experiment with learned neural evaluators and Bellman regressions, the method matches independent policies' final performance using two modules instead of three and substantially reduces forgetting relative to ungated sharing.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.