acceptodds
Under review as a conference paper at ICLR 2027

Every Bit Counts: Resource-Efficient Learning for Endangered Language Translation

Abstract

Despite rapid advances in large language models (LLMs), endangered languages remain severely underserved due to their limited representation in pretraining corpora and the scarcity of standardized linguistic resources. This challenge is particularly acute for machine translation, where existing approaches typically address resource integration, data augmentation, and learning optimization in isolation, lacking a systematic framework that connects them under extreme resource scarcity. To the best of our knowledge, we present the first unified framework that connects these three stages through model confidence as a shared signal, using Manchu translation as a case study. Rather than applying confidence uniformly, we adapt its role to the objective of each stage. First, Dictionary-guided Knowledge Correction (DKC) leverages bilingual lexical resources to correct high-confidence errors inherited from pretraining. Second, Reverse Lexical Mapping Augmentation (RLMA) expands limited supervision by promoting diversity across confidence levels. Third, Confidence-Aware Progressive Selection (CAPS) prioritizes low-confidence examples that remain insufficiently learned and progressively introduces them according to the evolving model state. On Manchu translation, our framework improves BLEU from 16.40 to 26.38. The results demonstrate the complementary roles of confidence in correcting unreliable knowledge, diversifying augmented supervision, and prioritizing unresolved examples, providing a unified approach to exploiting scarce linguistic resources for endangered language translation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.