DocGRPO: Adaptive Adequacy and Fluency Preference Optimization for LLM-Based Document-Level Machine Translation
Abstract
Large language models (LLMs) have shown strong potential for machine translation, yet preserving source information while maintaining coherent discourse remains challenging in document-level machine translation (DocMT). Our preliminary analysis shows that word-alignment-based coverage declines as source documents grow longer, whereas fluency improves and human quality judgments follow a non-monotonic pattern. Coverage alone, however, cannot distinguish genuine omissions from contextually appropriate ellipsis and implicit references, motivating the joint consideration of adequacy and fluency. We propose DocGRPO (ynamic bjective oordination with roup elative olicy ptimization), a reinforcement-learning framework that combines an alignment-based coverage reward for source-content preservation with a discourse-aware fluency reward. To further alleviate the tension between adequacy and fluency, DocGRPO adapts reward weights to training progress and source-document length while combining group-relative and reference-relative advantages to guide policy updates. Experiments across five translation directions and two model scales show improvements over strong document-level baselines. Human evaluation further confirms better overall translation quality, while ablations support the benefits of complementary rewards and adaptive weighting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.