acceptodds
Under review as a conference paper at ICLR 2027

DTSA: Dual-Timescale Stability and Adaptation for Robust Offline Reinforcement Learning with Spiking Neural Networks

Abstract

Combining offline reinforcement learning with spiking neural networks exposes a distributional instability that a state–action mismatch account does not capture: fixed data support interacts with time-varying neuronal dynamics. We formalize this as a dual-timescale shift: slow offline optimization drifts spike-conditioned action relations away from the supported structure, while deployment perturbations move a frozen policy into neural-dynamic regimes not covered by training. Treating the spike response induced by the decision history as an explicit condition variable, we propose DTSA, whose two mechanisms operate on one shared state–spike condition space: Conditional Structure Preservation (CSP) stabilizes spike-conditioned local action structure during training, and Adaptive Episodic Memory (AEM) corrects actions within an episode through confidence-gated, magnitude-limited fusion with all parameters frozen, functionally inspired by the stability–plasticity principles of biological learning. Both mechanisms are analyzed under a unified dual-timescale error decomposition. Experiments show that DTSA consistently improves the average normalized return across multiple benchmark tasks, by points on Adroit over the matched spiking backbone, and exhibits better overall performance retention and distribution-shift robustness under observation noise, sensor dropout, and environment-dynamics shifts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.