acceptodds
Under review as a conference paper at ICLR 2027

Reset Cost Does Not Determine Reset Value The Role of Post Reset Update Budgets

Abstract

Neuron resets aim to restore plasticity in deep reinforcement learning, and recent work frames the choice of which units to reset as a cost–benefit problem with an estimable immediate cost. Does that cost predict a reset's long-run value? From common SAC checkpoints we measure, for matched reset interventions, the immediate change in the critic's TD loss and the returns that follow. An operator-aligned sensitivity score ranks the held-out loss change of random masks accurately (Spearman ), where a signed first-order prediction does not. This local accuracy does not yield a reliable ordering of long-run value: within matched checkpoint–pressure regimes, pre-registered tests detect no monotonic cost–value relation that is stable across tasks and update pressures. Across six DeepMind Control tasks, repeatedly resetting the highest-scoring units is substantially less harmful under high update pressure than under low pressure on every task, despite several times the cumulative cost. A history recovery-budget factorial on two tasks separates the pre-reset training history from the post-reset update budget: in an algebraic decomposition the recovery budget accounts for of the pooled high-versus-low-pressure difference, and training history contributes a smaller, task-dependent share. Immediate cost is therefore measurable but insufficient to decide whether a reset is worthwhile; its value depends on the learning that follows. A cost estimator settles which reset is cheap, not which one is worth making.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.