acceptodds
Under review as a conference paper at ICLR 2027

MORPHOT: Evaluating Morphologically Rich Indic Languages with Hierarchical Unbalanced Optimal Transport

Abstract

Embedding-based translation metrics correlate strongly with human judgments on English benchmarks, yet the same metrics fail systematically on morphologically rich languages (MRLs) for three structural reasons. Such metrics stay blind to fine-grained inflection errors, conserve mass between hypothesis and reference and therefore conflate omissions with hallucinations, and reward reference mimicry without verifying fidelity to the source. We introduce MORPHOT, a hierarchical unbalanced optimal transport (UOT) metric that answers each failure mode with one architectural component. MORPHOT represents a sentence as a discrete measure over morphological units. An inner UOT over affix lattices produces a morphology cost sensitive to inflectional differences, an outer UOT decomposes into a matched discrepancy together with separate omission and hallucination residuals, and a source-anchored adequacy term guards against reference-mimicking drift. A monotone softplus-parameterized head then aggregates the components under a formal guarantee that no training run inverts the sign of a penalty. With  M trainable parameters atop a frozen multilingual encoder, MORPHOT attains the highest pooled Pearson, Spearman, and Kendall correlation with human quality ratings on Indic MRLs among all baselines evaluated, improving Pearson by over BERTScore and by over BLEURT-20, which carries more trainable parameters. The component decomposition additionally classifies Multidimensional Quality Metrics (MQM) error categories at macro-F1 against for the strongest scalar baseline.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.