Reward-Guided Noise-Level Trajectories and Task-Aware Evaluation for Image Editing
Abstract
Instruction-guided diffusion and flow-based image editing requires selecting a noise-level trajectory whose effectiveness can vary with the source image and instruction. Selecting such trajectories requires a reliable editing reward, yet the evidence needed to assess quality varies across tasks despite the shared goals of faithful editing, content preservation, and visual integrity. We address these two challenges, input-adaptive trajectory selection and task-aware reward construction. Amortized Trajectory Selection (ATS) learns a lightweight policy optimized for a chosen editing reward to predict a complete trajectory from each source–instruction pair before execution, without test-time optimization or online reward queries. To our knowledge, ATS is the first image-editing method to learn such direct per-input trajectory prediction without test-time optimization. Geometry-Normalized Trajectory Selection (GeoTS) complements ATS by making the sampling trajectory itself a structured source of candidate variation without training a selector or iteratively optimizing individual trajectories. To provide a reliable reward across editing tasks, we introduce EditCal, which selects and combines complementary signals from existing evaluation models according to the task and criterion, calibrates them to shared human-anchored ordinal semantics, and integrates them into an overall reward with criterion-specific failure diagnostics. On FLUX.1-Kontext and FlowEdit, ATS improves mean reward over native schedules across all four evaluated objectives and retains its gains at unseen sampling-step counts without retraining. Across three public human-judgment benchmarks, EditCal improves evaluation agreement and failure detection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.