acceptodds
Under review as a conference paper at ICLR 2027

The Fate of a Failed Training Run: Recovery Dynamics and Replay-Buffer Pathology in Soft Actor-Critic

Abstract

A failed reinforcement-learning run is usually reduced to one observation: its final score. We instead save the failed state and branch it repeatedly, turning post-failure recovery into a distribution over stochastic continuations. The resulting 720-run study covers 20 failed Soft Actor–Critic checkpoint branch points on two Meta-World tasks. Recovery is a multi-stage process: a first successful episode (a spark), optional crossing of a repeated-success marker (consolidation), and a final solved or unsolved evaluation; a consolidated failure is a relapse. No unsparked continuation solves (0/313), whereas 111/159 continuations that reach 50 successes solve. Critic-loss excursions are not terminal: 94/120 solved continuations cross the operational instability marker, and in 49/90 solved continuations containing both events the critic crosses first. Actor–critic reset raises the equal-checkpoint consolidation rate by 0.160 (95% checkpoint-blocked bootstrap interval ), although its solve-rate gain is less certain and heterogeneous across checkpoints. Replay data supplies a sharper intervention. On checkpoint s18, reset with the failed run's buffer yields 0/10 solves; clearing that buffer yields 4/10; and replacing it with two compatible zero-success foreign buffers yields 9/10 and 7/10. The matched donor comparison shows that successful donor episodes and post-reset donor history are not necessary for rescue in this setting, while leaving the enabling property of the replacement data open. Finally, a checkpoint-aware stopping frontier quantifies the cost of acting on delayed recovery: an illustrative k no-spark rule saves 24.6% of post-branch compute while sacrificing 7/120 eventual solves. Replicated branches therefore expose both the dynamics of recovery and a data-state obstruction that a parameter reset alone can preserve.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.