acceptodds
Under review as a conference paper at ICLR 2027

TransVis: Benchmarking Real-World In-Image Machine Translation at Scale and Fine Granularity

Abstract

In-image machine translation (IIMT) aims to translate the text inside an image while seamlessly preserving the original visual layout. Existing benchmarks either stop at textual translation or assess generated images with image-level metrics; even multi-aspect protocols do not trace each source text block through discovery, translation, text rendering and visual preservation. Consequently, omissions, mistranslations and visual damage remain difficult to attribute. We therefore present TransVis-Bench, a large-scale, fine-grained benchmark for real-world Chinese-to-English IIMT. The benchmark covers major categories and sub-categories of text–visual relations, comprising images and source text blocks. We design a source-block-anchored evaluation protocol that decouples model behaviour into four dimensions: translation coverage, translation quality, rendering accuracy and visual fidelity. An evaluation of foundation image-generation models shows that a single aggregate score hides markedly different combinations of failure in text translation, text rendering and visual preservation. Controlled experiments further show that explicitly injecting linguistic information raises the translation score of an open-source backbone from to , indicating that acquiring linguistic information remains a major limitation of current end-to-end translate-and-render models. TransVis-Bench provides a comprehensive testbed for real-world IIMT and supports auditable, fine-grained analysis of generated results.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.