acceptodds
Under review as a conference paper at ICLR 2027

Never Stop Rewiring: Safe Continual Reinforcement Learning on Non-Resetting Graphs

Abstract

A network agent that is to remain autonomous over a long horizon must keep learning as failure processes arrive, vanish and return, and it must do so without ever breaking the network, because on a graph that is never reset a wrong repair rewrites the topology that every later failure acts on. Continual RL supplies the first requirement and safe RL the second, but each is studied without the other: continual RL on benchmarks that reset between tasks, safe RL in episodes that undo a violation. We argue that lifelong network repair needs their combination, and study it as task-free safe continual RL on a non-resetting graph: a non-stationary constrained MDP with a service reward, a connectivity safety cost, no mode labels, and a budget-B unrecoverable event, a violation that no B edge additions can repair. Across Erdos–Rényi, Watts–Strogatz and Barabási–Albert ˝ streams with recurring and unseen failure modes, we obtain three findings. Reward forgetting, the failure continual RL is built to prevent, is near zero for every learner; what ends autonomy is an absorbing violation state into which the service reward leads the agent and out of which later learning does not climb. Continual learning alone makes this worse: PPO is absorbed in 17 of 30 lifetimes, policy consolidation in 26, because the cascade stores the unsafe behaviour; a Lagrangian constraint alone only delays it (12 of 30). Their combination, constrained policy consolidation, which prices the constraint on the visible policy of a multi-timescale cascade, keeps rewiring for 29 of 30 lifetimes with the lowest violation rate in every graph family and a violation rate that falls on every return to a mode, without task boundaries. For a network to become autonomous over its whole lifetime, its agent must therefore never stop learning and never stop being constrained: the reward it can relearn, the safe behaviour it must carry forward.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.