TextReVision: Benchmarking Image Text Editing towards Intent Fidelity with Gate-and-Craft Evaluation
Abstract
Reliable image text editing requires satisfying textual, spatial, stylistic, and semantic constraints while preserving non-target content. Existing benchmarks inadequately cover complex editing intents, while text recognition and image similarity alone cannot establish whether an instruction has been fulfilled. We present TextReVision, a bilingual, scenario-driven benchmark comprising 1,200 human- curated samples across 16 scenarios and five capability families: basic editing, layout control, style design, semantic reasoning, and multi-reference composition. Sample-specific rubrics provide 6,426 verifiable criteria across eight evaluation dimensions. Furthermore, we propose Gate-and-Craft scoring, in which essential task completion bounds the attainable score and scales the contribution of execution quality. The scoring rule achieves a concordance correlation coefficient of 0.935 against human reference scores. Across 18 models and 21 configurations, the leading scenario-balanced scores reach 89.5/89.3 on Chinese/English samples, yet semantic reasoning remains a major weakness. SenseNova-U1.5-8B-MOT’s official thinking mode raises its reasoning scores from 34.7/31.8 to 71.7/70.1, revealing that reasoning over user intent and translating it into actionable edits are critical. TextReVision advances evaluation toward verifiable, intent-grounded image text editing.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.