acceptodds
Under review as a conference paper at ICLR 2027

Can We Perform Online RL for Image Editing without Editing Rewards?

Abstract

Reinforcement learning (RL) enables direct preference optimization for image editing through editing-specific rewards, which remain less developed due to costly `edit instruction, source & edited image' triplet supervision and complex task-dependent calibration. In contrast, text-to-image (T2I) generation benefits from a mature and diverse reward ecosystem spanning semantic alignment, aesthetics, realism, glyph rendering, and other visual preferences. Extending this ecosystem to image editing would substantially broaden the range of visual preferences for RL-based optimization, raising the natural question: Can We Perform Image Editing RL without Editing Rewards? In this paper, we argue that the standard image editing dimensions have potential to be mapped to the T2I reward space: image quality can transfer directly, prompt following can be aligned through a description of the desired visual state, and reference consistency admits a coarse semantic conversion by encoding the source content to preserve. However, building such lever is non-trivial. Editing instructions specify relative changes, whereas T2I rewards require self-contained target descriptions; moreover, semantically correct captions to humans may not induce the frozen reward to rank successful edits above failures. We therefore introduce Lever-Edit, a two-stage reward-transfer framework. Stage 1 learns a textual reward interface that produces both semantically faithful and reward-aligned descriptions of desired post-edit states with a pluggable target-state adapter, change-aware source–target contrast, two-level preference calibration, and regularizations. Stage 2 freezes this interface and optimizes editing policy solely with the transferred T2I reward. Experiments show competitive editing alignment and source preservation against editing-reward-based fine-tuning, while outperforming intuitive transfer baselines {across various benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.