acceptodds
Under review as a conference paper at ICLR 2027

Beyond the Spectral Bulk: Spike Acceleration for Critic Learning in Deep Reinforcement Learning

Abstract

Value-function approximation can destabilize deep reinforcement learning because critic errors affect both bootstrapped targets and subsequent policy updates.Existing methods mainly stabilize learning targets, reduce estimation error, or constrain critic parameters, but do not directly promote weak task-related directionsthat remain unresolved within the spectral bulk. We propose Spike Spectral Acceleration Regularization (SSAR), which selectively amplifies critic singular modesnear the estimated upper bulk edge and removes the additional pressure after sufficient spectral separation.SSAR is motivated by the Baik–Ben Arous–P´ ech ´e (BBP) transition, under which an emerging low-rank signal becomes increasingly aligned with an empirical singular mode as it approaches and crosses the separation threshold. For a fixedtarget critic update, we decompose the SSAR-induced Bellman-loss change intoa task-aligned descent term and a remainder, and characterize their scaling acrossthe BBP transition. Under a rank-one spiked model, the two terms have the sameasymptotic order in the fixed-subcritical regime, providing no sign guarantee. Atnon-degenerate criticality and in the fixed-supercritical regime, the aligned termdominates, making the additional SSAR step loss-decreasing with probabilitytending to one.This analysis shows that spectral proximity determines where SSAR acts, whilethe emergence of task alignment determines when the additional update becomesbeneficial. Controlled experiments verify the predicted finite-width scaling andloss-sign transition. We further evaluate SSAR with Soft Actor-Critic on sixcontinuous-control tasks from the DeepMind Control Suite against vanilla SACand matched L2 critic regularization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.