Haidass: Bilingual Small Language Models Based on Data Curation and Language Scheduling
Abstract
Small language models (SLMs) offer clear advantages in efficiency and deployment, especially for on-device bilingual applications, but at the sub-200M scale, they remain English-centric: robust bilingual SLMs are scarce. Introducing a second language at this scale forces both languages to share a narrow parameter budget. To inform the design of such a bilingual SLM, we study how Chinese-English pretraining corpora interact. We quantify the cross-lingual interference between the two languages and how it depends on the mixing ratio, and compare several data strategies. Guided by this, we curate a balanced Chinese-English corpus and train Haidass-143M on 391.1B tokens, the first bilingual SLM at this scale, with a multi-stage real-to-synthetic data curation and language scheduling. It ranks 5th on the Open SLM Leaderboard. Building on this model, we post-train Haidass-Translate-143M. With direction-specific beam decoding, it achieves 30.36 BLEU and 21.40 chrF++ on FLORES-200 (enzh), outperforming the dedicated NLLB-200-distilled-600M by 7.92 BLEU and 4.66 chrF++, and essentially matches Qwen3-0.6B (30.94 BLEU, 21.10 chrF++) despite roughly 4 fewer parameters. This work provides empirical guidance for the training recipe design of SLMs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.