acceptodds
Under review as a conference paper at ICLR 2027

Breaking Babel: A Self-Evolving Multi-Agent System for Long-Form Subtitle Translation

Abstract

Long-form subtitle translation presents challenges beyond those of conventional machine translation: a single episode may contain hundreds of sentences whose meanings depend on long-range discourse and cultural context spanning the entire episode or even series, while translation quality also requires maintaining consistent terminology and style throughout the series. Existing approaches remain limited. Single-LLM methods operate at the sentence level, lacking long-context understanding and consistent terminology across episodes. Multi-agent methods often rely on static workflows that fail to adapt to scene complexity. In addition, both paradigms ignore the production context, e.g., genre. We propose SMART, a Self-evolving Multi-Agent system for long-foRm subtitle Translation. SMART operates in two stages: during test-time training, it builds a persistent series-level memory and iteratively translates a subset of sentences through a dynamic graph router and a Mixture-of-Agents translation layer, equipped with tool-calling modules for terminology verification, subtitle constraint validation, and contextual retrieval; a judge-refiner loop scores each candidate translation and back-propagates textual critiques that refine agent prompts and the routing policy, without retraining the underlying LLMs. During test-time inference, this evolved configuration translates the remaining sentences of the series. To evaluate long-form subtitle translation at scale, we introduce Subtitle Arena, covering 14 genres, 2–198 episodes per series, production years from 1959 to 2023, and 15 target locales, together with SubMQM, a subtitle-adapted MQM framework with seven dimensions and 19 error categories. SMART achieves the best overall MQM score in all 15 Subtitle Arena directions, reducing the average penalty by 6.9% over the strongest competing agent system. On a public benchmark, MuSC, SMART obtains the best model result across all 4 language pairs. SMART also achieves the best result in human evaluation with an overall score of 4.50/5.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.