acceptodds
Under review as a conference paper at ICLR 2027

Quality- and Difficulty-Aware Data Selection for Multilingual Extraction

Abstract

Training models on translated data is widely used to improve multilingual capabilities, particularly for low-resource languages where data scarcity remains a longstanding challenge. However, translated data can contain noise, making effective data selection crucial. While data quality can help identify noisy instances, it does not indicate how informative an instance is for learning. We therefore propose selecting translated data by jointly considering data quality and example difficulty. We study this approach in extraction tasks, where predictions are made over words or phrases and must be precisely grounded in the input text. Such tasks are particularly valuable targets for translation, since fine-grained annotation is costly to produce in new languages. However, they also pose a particular challenge: label annotations may become invalid after translation even when sentence-level translation quality is high. We show that machine translation metrics fail to capture such errors and complement them with label transfer accuracy, assessed using an LLM-as-a-judge. We evaluate these selection strategies by fine-tuning an LLM on English extraction tasks and their translations into 20 languages. Our results show that selection based on quality alone generally outperforms random selection, while prioritizing high-quality instances within a target difficulty range further outperforms quality-based selection alone on 7 of 8 evaluation datasets. Finally, we show that fine-tuning on translated data remains beneficial even for LLMs with strong multilingual abilities, outperforming few-shot prompting on 6 of 8 datasets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.