BeyondAdeEval: Machine Translation Evaluation Beyond Semantic Adequacy
Abstract
Machine translation metrics and LLM-based evaluators have made substantial progress in translation evaluation. However, existing methods such as COMET primarily focus on semantic adequacy and offer limited interpretability. Evaluating overall translation quality beyond semantic adequacy remains challenging. We present BeyondAdeEval (BAE), a reference-free and interpretable method for evaluating translation quality from semantic adequacy to finer-grained aspects of translation quality. We design , a principled, human-aligned hierarchical rubric that formalizes translation quality, guides data construction, and enables our models to learn structured, rubric-grounded rationales. Building on this rubric, we introduce (i) , a preference-graded translation benchmark constructed through controlled modification of high-quality translations, with 6,960 training samples and a human-annotated test set supporting pointwise scoring and pairwise ranking; and (ii) , a family of evaluation models ranging from 4B to 9B parameters that produce human-aligned, multidimensional, and explainable translation evaluations. On BAE-BENCH, our models outperform all evaluated baselines, achieving relative improvements of 28% and 24% in Spearman and Pearson correlation, respectively, over the strongest learned metrics. We also achieve state-of-the-art performance on T-index (Human) and demonstrate zero-shot cross-lingual generalization on MQM24. We will release our code and data upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.