ReflowText: Beyond Local Text Replacement for Coherent Visual Content Revision
Abstract
Image editing models have become effective at replacing short words and phrases, but real updates often involve long text, source-to-target length changes, and multiple jointly edited regions. Beyond accurate glyphs, models must adapt line breaks, text scale, occupied regions, and surrounding layout while preserving style and non-edited content. Existing benchmarks do not systematically isolate these challenges. We introduce , an English-to-English, replacement-only benchmark with 2,002 full images and 9,197 regions across four categories and 15 subcategories. It controls three coupled difficulty axes—long-form content, multi-region replacement, and length mismatch—and combines OCR metrics with a fine-tued VLM judge, , to evaluate text fidelity, font style, layout and reflow, and background preservation. ReflowJudge reaches 79.1% agreement with human annotations. We further propose , a plug-and-play layout-first strategy in which a fine-tuned VLM predicts line-level layouts, a deterministic renderer creates spatially aligned glyph conditions, and the image editor restores visual appearance. Across two open-weight backbones, GlyphPlan consistently improves aggregate quality, especially font style and layout. Together, ReflowText and GlyphPlan advance reliable long-form, multi-region visual text editing with coherent layout adaptation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.