acceptodds
Under review as a conference paper at ICLR 2027

Scored by Its Own Reward: Evaluator-Dependent Dense-Reward Speed-Ups in Diffusion RL

Abstract

The measured advantage of a graded reward can depend on the evaluator used to define success. We compare a binary CLIP-threshold reward with a dense clipped-linear CLIP reward during DDPO fine-tuning of Stable Diffusion 1.5 on 14 compositional prompts. Two campaigns record CLIP scores; campaign B additionally scores each rollout with OWLv2 and BLIP-VQA, neither used for optimization. Hard-setting thresholds target a nominal 10% base-policy success rate. The dense reward reaches the CLIP success-rate target of 0.5 earlier by 8.45 and 7.19 epochs in the two campaigns (26.8% and 22.6% fewer epochs). In campaign B, on the 13 prompts with identifiable thresholds for all three evaluators, success changes from the initial block to the final ten-block mean from .104 to .914/.901 under CLIP, from .125 to .134/.146 under OWLv2, and from .121 to .182/.193 under BLIP-VQA (dense/binary). On these prompts, the dense-minus-binary success-rate AUC effects are +.0511, −.0101, and −.0006, respectively. The two CLIP-minus-external contrasts are +.0612 and +.0517, with simultaneous 95% confidence intervals [.0359, .0865] and [.0293, .0741]; complementary sign-flip tests give . Under these evaluators, the dense-reward speed-up is largely specific to the optimized CLIP score.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.