acceptodds
Under review as a conference paper at ICLR 2027

Blame Errant Words, Not the Canvas: Defect-Aware Spatial Credit Assignment for Visual Text Generation

Abstract

Recent reinforcement learning approaches have improved visual text generation, yet generated images still suffer from incorrect wording and malformed glyphs. We identify two coupled limitations: recognition-based rewards often overlook visible glyph defects, while whole-image scalar feedback obscures the uneven quality of individual text regions. These limitations motivate our central principle: *Blame errant words, not the canvas*, to effectively preserve local quality distinctions from evaluation to policy optimization. First, we introduce TextDART, a defect-aware region-targeted reward model that localizes individual text regions and independently evaluates their semantic correctness and glyph quality, providing fine-grained, region-precise feedback. Building on TextDART, we propose SCANNER, which turns these regional judgments into spatially differentiated policy feedback. It normalizes rewards for each text target across sampled images and casts the resulting regional advantages onto the corresponding latent tokens. Extensive experiments on existing benchmarks and our newly established PolyText-Bench demonstrate that our approach achieves state-of-the-art performance in text accuracy and glyph fidelity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.