acceptodds
Under review as a conference paper at ICLR 2027

MoRE-OPD: Mixture of Reward Experts with On-Policy Distillation for Image Editing

Abstract

Reliable evaluation is critical for image editing, supporting data curation, benchmarking, and reinforcement learning. However, diverse editing tasks require distinct criteria, leading to potential criterion entanglement in a shared evaluator. To address this challenge, we propose MoRE-OPD, a reward modeling framework that combines a Mixture of Reward Experts with privileged On-Policy Distillation. MoRE-OPD first learns category-specialized reward experts under task-specific evaluation protocols and then distills their heterogeneous evaluation capabilities into a compact universal evaluator. During OPD, specialized experts additionally access oracle-generated rationales as privileged context, enabling more effective supervision of student-generated predictions. Furthermore, we construct PointReward-Bench for reliable pointwise evaluation of image editing reward models, and apply MoRE-OPD to curate approximately 3.8 million high-quality samples for downstream optimization. Extensive experiments demonstrate that MoRE-OPD consistently outperforms state-of-the-art open-source image editing reward models across pointwise evaluation, preference alignment, and downstream reinforcement learning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.