acceptodds
Under review as a conference paper at ICLR 2027

SpikeARC: Spike-Guided Routing and Selective Policy Consolidation for Continual Reinforcement Learning

Abstract

Reinforcement learning can produce high-performance control policies, but real-world deployment demands more than performance in fixed environments. As environment dynamics and task requirements change, agents must keep learning, which can disrupt acquired yet still useful behaviors. Continual reinforcement learning therefore faces a fundamental challenge: when adaptation is required, it is not immediately clear whether a behavioral correction should remain transient or eventually become part of the long-term policy. We introduce SpikeARC, a framework that organizes continual policy adaptation through state-dependent action corrections between complementary Stable and Fast pathways. The Stable branch provides the primary action, while recurrent spiking Fast and Fusion modules generate and gate state-dependent corrections. Spike-guided routing assigns sample-wise replay responsibility, dual behavior-space constraints regulate behavioral drift, and corrections that remain consistent are progressively consolidated into the Stable branch. We systematically evaluate SpikeARC through external baseline comparisons, cross-environment controlled experiments and mechanism ablations, and structured experiments with simulated quadruped robots. Across external continual RL evaluations on CARL Walker and Quadruped, SpikeARC maintains competitive adaptation while reducing forgetting relative to hidden-context continual RL baselines. In a controlled five-environment study with a PPO backbone, it retains 90.8% of the adaptation AUC of a fully shared policy while reducing forgetting by 36.2% and long-term-policy parameter drift by 35.3% on average across environments. Ablations and structured quadruped experiments further characterize the adaptation–retention trade-offs induced by routing and consolidation. Overall, SpikeARC does not require a new behavioral correction to be classified as transient or permanent at first encounter; instead, its role is resolved progressively through continued interaction while mature behavior remains protected.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.