acceptodds
Under review as a conference paper at ICLR 2027

ConServe: Context-Preserving and Compute-Conserving Text Image Machine Translation

Abstract

Text Image Machine Translation (TIMT) requires models to recognize embedded text and use visual context to resolve translation ambiguities. However, existing approaches leave unclear which visual evidence supports translation decisions and how much context is necessary. We introduce Translation Context Lens (TCL), a causal diagnostic framework based on controlled visual interventions. TCL reveals that visual evidence for translation is distributed yet redundant, with different regions providing overlapping support. This finding motivates ConServe, a framework for preserving translation quality under visual context compression. Spatially balanced token sampling distributes the background budget across regions and selects representative tokens locally, preserving all visual tokens in target text regions. Privileged Context On-Policy Self-Distillation (PC-OPSD) addresses shifts in translation preferences caused by compression. The teacher uses full visual context, correct interpretations, and explanations of incorrect alternatives to supervise predictions on prefixes generated by the student. We also introduce SC-TIMT, a multilingual benchmark for screen translation that depends on visual context. Experiments show that our 9B model outperforms larger multimodal models on SC-TIMT and remains competitive on MCiT with substantially fewer visual tokens. These results suggest that efficient TIMT depends on preserving distributed contextual evidence rather than retaining the entire visual input.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.