Quality- and Difficulty-Aware Data Selection for Multilingual Extraction
Abstract
Training models on translated data is widely used to improve multilingual capabilities, particularly for low-resource languages where data scarcity remains a longstanding challenge. However, translated data can contain noise, making effective data selection crucial. While data quality can help identify noisy instances, it does not indicate how informative an instance is for learning. We therefore propose selecting translated data by jointly considering data quality and example difficulty. We study this approach in extraction tasks, where predictions are made over words or phrases and must be precisely grounded in the input text. Such tasks are particularly valuable targets for translation, since fine-grained annotation is costly to produce in new languages. However, they also pose a particular challenge: label annotations may become invalid after translation even when sentence-level translation quality is high. We show that machine translation metrics fail to capture such errors and complement them with label transfer accuracy, assessed using an LLM-as-a-judge. We evaluate these selection strategies by fine-tuning an LLM on English extraction tasks and their translations into 20 languages. Our results show that selection based on quality alone generally outperforms random selection, while prioritizing high-quality instances within a target difficulty range further outperforms quality-based selection alone on 7 of 8 evaluation datasets. Finally, we show that fine-tuning on translated data remains beneficial even for LLMs with strong multilingual abilities, outperforming few-shot prompting on 6 of 8 datasets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.