Search Finds Blind Spots: Stress-Testing Edit Reward Models Under Optimization Pressure
Abstract
Image editing reward models guide candidate selection and iterative refinement by scoring images. While evaluations measure agreement with human ratings, they provide limited evidence about reliability under sustained optimization. To address this gap, we present **EditStress-Bench**, a benchmark for evaluating editing rewards under search pressure. It comprises 300 human-object editing tasks, matched search controls, and independent human assessment of quality and editing defects. Experiments evaluate twelve reward settings from six reward families using Best-of- selection and iterative refinement. The strongest static reward becomes the weakest Best-of- selector at a budget of 64 candidates, revealing a mismatch between *static agreement* and *search-time quality*. Rewards with source-image access remain 10.6 Overall Human Quality (OHQ) points below the human oracle on average. Motivated by these findings, we introduce **factor-guided reward patching**, which profiles defect sensitivity, predicts selected failures, and targets weak factors during training. Targeted updates improve OHQ by 1.6 to 2.1 points, compared with 0.3 to 0.4 for equal-budget random updates. Repeated patching reduces monitored failures, while a broader audit identifies persistent appearance drift. Code, evaluation data, and human annotations will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.