acceptodds
Under review as a conference paper at ICLR 2027

Continual Learning with Success-Conditioned Self-Distillation

Abstract

Reinforcement learning with binary verifiable rewards provides only one bit of feedback per rollout, making credit assignment across thousands of tokens difficult. We argue that the success-conditioned policy provides a natural dense learning target for both single-task and continual learning by restricting the reference model to its own successful responses, which is the least disruptive way to succeed. Computing this target, however, requires a token-level value function that many verifiable-reward methods avoid learning. Prompting a frozen model with one of its own verified successes offers a full-vocabulary surrogate without an explicit critic. Yet this surrogate conditions on a particular successful trajectory rather than on success itself, introducing trajectory-specific nuisance information. We propose Success-Conditioned Self-Distillation (SCSD), which addresses this mismatch by searching for a textual transformation of a successful rollout that brings the prompted policy closer to the success-conditioned target. The resulting teacher is then distilled into a standalone student policy through on-policy distillation on the student's own rollouts. Across multiple tasks, SCSD achieves state-of-the-art performance in both single-task and continual learning settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.