What Makes a Face Edit Reliable? Benchmarking Reliability in Text-Guided Face Editing
Abstract
Text-guided face editing requires fulfilling editing instructions while preserving the subject’s identity. However, existing evaluation practices primarily emphasize edit effectiveness, leaving identity preservation insufficiently examined and potentially favoring models that follow instructions at the expense of identity. To address this gap, we introduce FaceEdit-Bench, a benchmark that evaluates identity preser- vation and edit effectiveness as complementary dimensions of editing reliability. FaceEdit-Bench comprises approximately 110K annotated face edits generated by 15 editing models across 20 instruction types. We assess identity preservation by aggregating scores from multiple face recognition systems. To evaluate edit effectiveness, we develop an instruction-conditioned agent that decomposes each instruction into verifiable requirements and assesses their fulfillment using tool-derived visual evidence. Our analysis reveals divergent model rankings across the two dimensions: stronger instruction fulfillment does not necessarily imply better identity preservation, highlighting the limitations of instruction fulfillment alone as a measure of editing reliability. We further introduce FaceEditEval, a vision-language model that consolidates supervision from the recognition systems and evaluation agent into a single evaluator, enabling efficient prediction of identity preservation and edit effectiveness as separate scores. FaceEditEval achieves higher correlations with the automated reference scores than existing quality assessment methods. By accounting for both the requested changes and the preservation of identity, FaceEdit-Bench and FaceEditEval enable more comprehensive comparisons of editing models and support research on reliable text-guided face editing.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.