FINDING THE PARALLELS: A PATH TO EVALUATING SUMMARIZATION FROM LOW-RESOURCE LANGUAGES
Abstract
Across benchmarks, large language models have demonstrated impressive summarization capabilities for high-resource languages. In contrast, summarization benchmarks are largely unavailable for low-resource languages. This scarcity poses a challenge for both evaluating LLM summarization capabilities and measuring the effectiveness of LLM-as-a-Judge approaches in these languages. However, previous work in machine translation has contributed parallel datasets that pair sentences in low-resource languages with higher-resource translations. We introduce PARAWEAVE, a method for automatically transforming these existing sentence-level parallel datasets into coherent document-level benchmarks while preserving their original cross-lingual alignment. We used PARAWEAVE to construct a summarization benchmark that spans ten language varieties from three families (Afro-Asiatic, Indo-Iranian, and Nilo-Saharan). We use the resulting benchmark to evaluate summarization faithfulness and LLM-as-a-Judge performance for summaries generated from documents in these languages. Pairing low-resource language documents with high-resource translations enables studying the effect of source language on both summarization quality and LLM-as-a-judge performance. These comparisons reveal gaps in model performance and demonstrate that even current frontier models struggle with faithful low-resource language summarization and evaluation in certain languages. For example, GPT-5.5's summaries of Nilo-Saharan documents are significantly less faithful than their English document counterparts. Native-speaker annotations validate the model judgments in our construction pipeline and judge faithfulness metrics. The proposed pipeline offers a general approach for repurposing existing sentence-level parallel resources for evaluation of document-level tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.