acceptodds
Under review as a conference paper at ICLR 2027

Reinforcement Continual Learning for One-step Image Generation Models

Abstract

We introduce Reinforcement Continual Learning (RCL), a post-training framework that converts evaluation feedback into image-level supervision through a generate–construct–select–update loop. For each current-model output, RCL constructs candidate alternatives, determines preference eligibility, selects a suitable target, and learns from it at the original noise input and conditioning through an independently specified matching loss. Distinguishing preference eligibility, source–target compatibility, and matching allows black-box evaluation to guide learning while preserving the inference path. We instantiate RCL on a pMF-H one-step ImageNet generator with retrieved targets and ImageReward or Qwen preferences. Under their respective preference criteria, ImageReward and Qwen scores continue to improve throughout post-training. In comparison with random target-selection policies, the ImageReward-guided variant forms the empirical Pareto frontier for ImageReward and the multi-representation Fr\'echet-distance ratio fdr6. In a single-run comparison with Direct ImageReward, RCL also yields a more favorable ImageReward–fdr6 trajectory. These results demonstrate the framework's feasibility and characterize the preference–distribution trade-offs of its retrieval-based instantiation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.