acceptodds
Under review as a conference paper at ICLR 2027

Search Finds Blind Spots: Stress-Testing Edit Reward Models Under Optimization Pressure

Abstract

Image editing reward models guide candidate selection and iterative refinement by scoring images. While evaluations measure agreement with human ratings, they provide limited evidence about reliability under sustained optimization. To address this gap, we present **EditStress-Bench**, a benchmark for evaluating editing rewards under search pressure. It comprises 300 human-object editing tasks, matched search controls, and independent human assessment of quality and editing defects. Experiments evaluate twelve reward settings from six reward families using Best-of- selection and iterative refinement. The strongest static reward becomes the weakest Best-of- selector at a budget of 64 candidates, revealing a mismatch between *static agreement* and *search-time quality*. Rewards with source-image access remain 10.6 Overall Human Quality (OHQ) points below the human oracle on average. Motivated by these findings, we introduce **factor-guided reward patching**, which profiles defect sensitivity, predicts selected failures, and targets weak factors during training. Targeted updates improve OHQ by 1.6 to 2.1 points, compared with 0.3 to 0.4 for equal-budget random updates. Repeated patching reduces monitored failures, while a broader audit identifies persistent appearance drift. Code, evaluation data, and human annotations will be released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.