Beyond Words Alone: Disentangling Distributed Evidence for Visually Grounded Cultural Translation
Abstract
Social media enables cross-lingual communication, yet the meaning of user-generated content (UGC) is often distributed across text, images, and cultural knowledge, making translation challenging. Existing evaluations typically treat images as optional and assume textual information is sufficient, obscuring whether errors stem from missing visual evidence or inadequate cultural adaptation. We introduce Culture-MMT, a diagnostic benchmark for distributed-evidence pragmatic reconstruction in social-media translation. It contains 498 curated, de-identified UGC instances across 20 domains, categorized into Symbol, Visual, and Hybrid by their required evidence. We also develop a human-aligned domain-specific judge, Judger, and a controlled evaluation framework that systematically varies the evidence available to translators. Across 17 model configurations, real images improve performance by +0.280 on Visual and +0.197 on Hybrid, but only +0.050 on Symbol, while random images provide no comparable gains. Human-authored descriptions further show that jointly dependent cases remain more challenging even when visual facts are explicit. These results characterize culturally effective social-media translation as meaning reconstruction from distributed evidence, highlighting visual grounding and cultural-pragmatic adaptation as distinct model bottlenecks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.