acceptodds
Under review as a conference paper at ICLR 2027

AffectEdit: Diagnosing and Fixing Evaluation for Affective Image Editing

Abstract

Image editors can now change the emotion an image conveys, yet it is unclear whether the metrics used to evaluate such affective edits agree with human judgment. To test them, we collect AffectEdit, a benchmark of 1,406 edits, each rated by about five annotators for whether it conveys the target emotion while preserving the scene. We find that the metrics in current use for affective editing fall short because they assess these two requirements separately. Emotion-only metrics can reward an edit that replaces the entire scene, scene-only metrics reward an edit that keeps the scene but leaves the emotion unchanged, and the one metric in current use that scores both adds the two terms, so either can mask a failure of the other. Replacing its addition with a multiplication, on the metric's own terms, makes it reject unchanged images and significantly raises its correlation with human ratings over the whole benchmark. We therefore propose PAEF (Perceived Affective Edit Fidelity), a parameter-free metric that multiplies an edit's emotion shift by its scene preservation, so that an edit scores highly only when both requirements hold. PAEF agrees with human ratings better than every metric in current use for affective edits, and holds that level on a held-out set of edits that informed none of its design choices. The strongest multimodal large language models (MLLMs), prompted or fine-tuned as judges, agree better still, but PAEF runs 3.7 to 61 times faster than an MLLM judge on the same GPU and needs no rated data. We release the benchmark, the ratings and our code (supplementary material).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.